modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
74 repositories
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
Гига Писарь: бесплатная локальная русская диктовка для macOS на GigaAM v3. Зажал правый ⌘, сказал, отпустил. Giga Pisar: offline Russian dictation for Mac
听记 (Tingji) — 本地会议录音转写与纪要,FunASR + LLM,数据不出本机。Local meeting transcription & minutes, runs entirely offline.
Распознавание русской речи в командной строке: GigaAM v3, ONNX int8, на процессоре, без интернета
Terminal voice-to-text TUI for Apple Silicon. Qwen3-ASR-1.7B or live-streaming Confucius4-R2T2 on the Apple GPU via MLX (mlx-speech). Fully local, no PyTorch.
VoxMinutes - free, local-first meeting assistant for Windows. Records system audio & mic together; real-time transcription, translation (13 languages) & AI summaries. Data never leaves your device.
Free, open-source offline dictation and transcription for macOS. Dictate into any app, transcribe calls and files on-device, and give Claude Code, Codex and Cursor ears via a local MCP server.
Lamitype (formerly HushType): free, privacy-first voice-to-text for macOS. Traditional-Chinese-first, memory-light. Runs Qwen3-ASR locally on Apple Silicon via MLX; optional cloud lane straight to your provider. 免費開源的 Mac 繁體中文語音輸入,本機執行 Qwen3-ASR。
Local-first dictation, voice notes & meeting memory for macOS — ⌘B anywhere, live transcript, on-device Gemma 4 / Parakeet models, zero cloud. MIT.
Talk. Ink. Push-to-talk dictation for macOS, 100% on-device. Pick your model: Qwen3-ASR, NVIDIA Nemotron or Voxtral, all via Apple MLX.
🎙️ The open-source AI Voice-to-Text & Universal Speech Studio. Real-time dictation, 48kHz Web Audio DSP noise cancellation, 2-Way Babel Live Translator with authentic Urdu/multilingual Neural TTS, Executive MoM PDF generator, Studio EQ mastering, and 100% offline audio tools. Powered by Gemini 2.0, GPT-4o & Claude 3.7.
Free Wispr Flow & Superwhisper alternative. Native macOS voice dictation & speech-to-text. Powered by Sber GigaAM v3 & Whisper. Instant direct input under cursor via ⌥+Space, push-to-talk, offline transcription of any media files.
Local voice dictation for Windows — what Hex is on macOS. Hold a hotkey, speak, release: the text lands at your cursor. Parakeet TDT v3, offline, ~0.2 s.
Voice dictation for the browser — free, private, MIT. Dictate into any web page, or use the pop-out to dictate for any app on your machine.
OpenAI-compatible speech-to-text server for nvidia/nemotron-3.5-asr-streaming-0.6b (NeMo). Runs on the DGX Spark / GB10.
CLI that turns a YouTube URL into a text transcript: yt-dlp for the audio, Whisper for the transcription, running 100% locally with GPU support.
本地实时课堂双语字幕 · 全本地英→中字幕,说话人一句定稿即出中文(ClassLive)
High-performance, local-first speech recognition (ASR) & synthesis (TTS) runtime with OpenAI-compatible APIs, optimized for Apple Silicon (MLX).
Real-time ASR WebUI on Apple Silicon (pure Rust + MLX): Qwen3-ASR transcription with auto language detection, mic/system-audio capture, AI polish/translate, meeting & content AI summaries, subtitle mode, domain terminology config, three-column UI, idle model unload
Sono is on-device dictation for macOS. Parakeet v3 + Apple Intelligence, nothing leaves your Mac.
Free, open-source, fully-local dictation for macOS. Hold a key, speak, release: clean text at your cursor. A Wispr Flow alternative that never touches the cloud.
🎬 面向内容创作者的 AI 字幕工作流助手,基于 Gemini 2.5 Pro。独创“音频理解 + 术语人工确认”两阶段工作流,从源头解决技术口播与专有名词识别翻车问题,支持一键导出 SRT/VTT/ASS。
Your voice, typed. Free, open-source dictation for Mac that turns speech into clean text in any app, entirely on-device.
Live speech-to-text streaming on Apple Silicon — Qwen3-ASR + Silero VAD + MLX
Transcribe any audio file to text locally on an Apple Silicon Mac — free, private, no cloud APIs. Powered by NVIDIA Parakeet + Apple MLX.
🎬 AI subtitle generator: convert video to SRT subtitles locally with NVIDIA NeMo Parakeet-TDT speech-to-text. GPU-accelerated, word-level timestamps, VAD, LLM correction — a fast offline Whisper alternative.
Dictado por voz para Windows. Push-to-talk, pulido con LLM, diccionario personal. Espanol e ingles mezclados.