Topic: speech-recognition

9,080 repositories

huggingface/transformers

πŸ€— Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

β˜… 166,888PythonForks 34,737

mozilla/DeepSpeech

DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.

β˜… 26,771C++Forks 4,073

m-bain/whisperX

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

β˜… 24,326PythonForks 2,454

modelscope/FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

β˜… 20,565PythonForks 2,054

leon-ai/leon

🧠 Leon is your open-source personal assistant.

β˜… 17,551TypeScriptForks 1,472

kaldi-asr/kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

β˜… 15,492ShellForks 5,352

alphacep/vosk-api

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

β˜… 15,158Jupyter NotebookForks 1,768

NVIDIA/DeepLearningExamples

State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.

β˜… 14,857Jupyter NotebookForks 3,407

abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

β˜… 12,967PythonForks 1,870

kmario23/deep-learning-drizzle

Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures!!

β˜… 12,958HTMLForks 2,982

PaddlePaddle/PaddleSpeech

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

β˜… 12,690PythonForks 1,955

QuentinFuxa/WhisperLiveKit

Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

β˜… 11,108PythonForks 1,143

openvinotoolkit/openvino

OpenVINOβ„’ is an open source toolkit for optimizing and deploying AI inference

β˜… 10,940C++Forks 3,423

espnet/espnet

End-to-End Speech Processing Toolkit

β˜… 9,976PythonForks 2,438

QwenAudio/SenseVoice

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

β˜… 9,424CForks 833

Uberi/speech_recognition

Speech recognition module for Python, supporting several engines and APIs, online and offline.

β˜… 8,994PythonForks 2,417

nl8590687/ASRT_SpeechRecognition

A Deep-Learning-Based Chinese Speech Recognition System εŸΊδΊŽζ·±εΊ¦ε­¦δΉ ηš„δΈ­ζ–‡θ―­ιŸ³θ―†εˆ«η³»η»Ÿ

β˜… 8,391PythonForks 1,891

Blaizzy/mlx-audio

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

β˜… 7,966PythonForks 736

TalAter/annyang

πŸ’¬ Speech recognition for your site

β˜… 6,818TypeScriptForks 1,047

flashlight/wav2letter

Facebook AI Research's Automatic Speech Recognition Toolkit

β˜… 6,437C++Forks 987

modelscope/FunClip

FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

β˜… 6,360PythonForks 757

wenet-e2e/wenet

Production First and Production Ready End-to-End Speech Recognition Toolkit

β˜… 5,243PythonForks 1,189

Picovoice/porcupine

On-device wake word detection powered by deep learning

β˜… 4,944PythonForks 581

yanshengjia/ml-road

Machine Learning and Agentic AI Resources, Practice and Research

β˜… 4,943PythonForks 1,721