Feature Request: Add SenseVoice + cam++ as an alternative ASR+diarization pipeline

#370 · closed · 0 comments

View on GitHub ↗

LauraGPT

whisper-diarization does a great job combining Whisper with pyannote for ASR + speaker diarization. I'd like to suggest an alternative pipeline using **FunASR**, which bundles both ASR and speaker diarization in a single toolkit: ### FunASR Pipeline - **SenseVoice** (ASR): 170x realtime, 234M params, 50+ languages, built-in VAD + punctuation - **cam++** (Speaker diarization): Lightweight 7.2M parameter speaker embedding model, state-of-the-art on CN-Celeb and VoxCeleb ### Advantages over Whisper + pyannote 1. **Much faster**: SenseVoice is non-autoregressive (170x vs ~13x realtime) 2. **Unified toolkit**: ASR, VAD, punctuation, and diarization all in one pip install 3. **Smaller models**: SenseVoice (234M) + cam++ (7.2M) vs Whisper large (1.5B) + pyannote 4. **Built-in features**: Emotion detection, audio events (laughter, applause) ### Quick example ```python pip install funasr from funasr import AutoModel # ASR with speaker diarization model = AutoModel( model="iic/SenseVoiceSmall", vad_model="fsmn-vad", spk_model="cam++", ) result = model.generate(input="meeting.wav") ``` ### Resources - FunASR: https://github.com/modelscope/FunASR (16.6K+ stars) - SenseVoice: https://github.com/FunAudioLLM/SenseVoice (8.3K+ stars) - cam++ paper: https://arxiv.org/abs/2303.00332 Would you consider adding this as an alternative pipeline option?

Comments