LauraGPT/funasr_stt
FunASR/SenseVoice microphone STT extension for oobabooga textgen
FunASR/SenseVoice microphone STT extension for oobabooga textgen
Agentic Development Environment based on OpenCode AI agent
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A high-throughput and memory-efficient inference and serving engine for LLMs
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
Powerful AI Client
A free, open source, and extensible speech-to-text application that works completely offline.
🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.
A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw
Voice-to-text with push-to-talk for Wayland compositors
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
Create polished demo videos without editing skills. Mac/Windows/Linux
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features
End-to-End Speech Processing Toolkit
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.
Bilibili视频转文字,一步到位,输入链接即可使用
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
直播间 AI 互动助手(快手+抖音):实时采集弹幕与主播语音,LLM 生成评论自动发送,帮你的直播间热场
🎤 开源免费的 AI 语音输入工具 — 基于FunASR,本地处理,为中文而生 | Free & open-source voice input tool with local FunASR
Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
Integrate cutting-edge LLM technology quickly and easily into your apps
Mock channel server CLI for OpenClaw testing
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.