Deep-unlearning/kernels-community
Kernel sources for https://huggingface.co/kernels-community
Open Source Audio MLE at ๐ค
Kernel sources for https://huggingface.co/kernels-community
Pure-PyTorch inference for CohereLabs/cohere-transcribe-03-2026 (2B Conformer + Transformer ASR, 14 languages).
Practical, Colab-friendly notebooks for fine-tuning and running audio AI models
The agent that grows with you
LLaSA: Scaling Train-time and Test-time Compute for LLaMA-based Speech Synthesis
Build local voice agents with open-source models
A ComfyUI custom node integration for multi-language High-quality Text-to-Speech and Voice Conversion nodes using multiple engines like RVC, ResembleAI's Chatterbox TTS, F5-TTS, Higgs Audio 2 and Microsoft VibeVoice with unlimited text length, SRT timing, Character support, Audio Analyzer, Silent Speech Analyzer, audio edit and more!!
A simple implementation for improving CosyVoice2 by GRPO method
Liquid Audio - Speech-to-Speech audio models by Liquid AI
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Fine-tuning & Reinforcement Learning for LLMs. ๐ฆฅ Train OpenAI gpt-oss, DeepSeek, Qwen, Llama, Gemma, TTS 2x faster with 70% less VRAM.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. ๐ฆ
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
A PyTorch native library for large model training
Fun-ASR is an end-to-end speech recognition large model launched by Tongyi Lab.
[ICASSP2025] Official code for VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Omnilingual ASR Open-Source Multilingual SpeechRecognition for 1600+ Languages
Train transformer language models with reinforcement learning.