Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models

Published 16 Jun 2025 in cs.CL, cs.AI, cs.SD, and eess.AS | (2506.13300v3)

Abstract: This paper presents Seewo's systems for both tracks of the Multilingual Conversational Speech LLM Challenge (MLC-SLM), addressing automatic speech recognition (ASR) and speaker diarization with ASR (SD-ASR). We introduce a multi-stage training pipeline that explicitly enhances reasoning and self-correction in speech LLMs for ASR. Our approach combines curriculum learning for progressive capability acquisition, Chain-of-Thought data augmentation to foster intermediate reflection, and Reinforcement Learning with Verifiable Rewards (RLVR) to further refine self-correction through reward-driven optimization. This approach achieves substantial improvements over the official challenge baselines. On the evaluation set, our best system attains a WER/CER of 11.57% for Track 1 and a tcpWER/tcpCER of 17.67% for Track 2. Comprehensive ablation studies demonstrate the effectiveness of each component under challenge constraints.