Feature: Add FunASR as a recognition backend (recognize_funasr)

#897 · open · 1 comments

View on GitHub ↗

LauraGPT

<!-- funasr-ops:accuracy-note-20260714 --> > [!NOTE] > **License and capability clarification (2026-07-14):** FunASR is a toolkit, not a single checkpoint. The [FunASR](https://github.com/modelscope/FunASR#license) and [SenseVoice](https://github.com/FunAudioLLM/SenseVoice#license) repository source code is MIT; model weights follow each model card. [SenseVoiceSmall](https://huggingface.co/FunAudioLLM/SenseVoiceSmall) supports Chinese, Cantonese, English, Japanese, and Korean, and its weights use the linked FunASR Model Open Source License Agreement. [Fun-ASR-Nano-2512](https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512) is Apache-2.0. Language coverage, punctuation, and performance depend on the selected model and runtime configuration. ## Feature Request The library currently supports Google, Sphinx, Vosk, Whisper, Faster Whisper, and OpenAI/Groq APIs. Proposing to add [FunASR](https://github.com/modelscope/FunASR) as an additional backend. **Why FunASR**: - **170× realtime on GPU** — SenseVoice is non-autoregressive, significantly faster than Whisper - **50+ languages** with built-in VAD, punctuation restoration, speaker diarization - **Works offline** — like Vosk and Whisper, fully local - **OpenAI-compatible API** — `funasr-server` at `/v1/audio/transcriptions` - **16K+ GitHub stars**, active development **Proposed API** (following existing patterns): ```python import speech_recognition as sr r = sr.Recognizer() with sr.AudioFile("audio.wav") as source: audio = r.record(source) # Local model text = r.recognize_funasr(audio, model="iic/SenseVoiceSmall") # Or via OpenAI-compatible server text = r.recognize_funasr(audio, api_url="http://localhost:8000") ``` **Integration is straightforward** — FunASR's Python API: ```python from funasr import AutoModel model = AutoModel(model="iic/SenseVoiceSmall") result = model.generate(input=audio_data) text = result[0]["text"] ``` Happy to submit a PR if there's interest. - [FunASR](https://github.com/modelscope/FunASR) — 16K+ stars - [SenseVoice](https://github.com/FunAudioLLM/SenseVoice) — 8K+ stars

Comments

LauraGPT

A first implementation for this request is open in #903. It adds `Recognizer.recognize_funasr(...)` as an optional local recognizer backed by FunASR/SenseVoice/Paraformer, with `SpeechRecognition[funasr]` documented as the optional install path. Current head `2dcc004a280a13e23f211acfc234c9ee6dfa37d7` is mergeable. Local validation covers the focused FunASR recognizer tests, the existing recognition smoke slice, Python compilation, RST checks, changed-file flake8, and a real CPU SenseVoice smoke. The remaining GitHub-side gate is maintainer approval of the fork workflow suites; there is no current code-owned failure visible on the PR.