Phd. Realtime-interaction.
xzf-thu.github.io
227 followers11 repositories
Repositories
Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing information and emotions respectively, while a fully streaming architecture eliminates latency at the fundamental level.
★ 2,283PythonForks 173
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
★ 1,146PythonForks 75
★ 0TypeScriptForks 0
The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.
★ 300PythonForks 24
★ 592PythonForks 33
Mini-Omni-Reasoner: a real-time speech reasoning framework that interleaves silent reasoning tokens with spoken response tokens (“thinking-in-speaking”), exploiting the LLM–audio throughput gap to keep speech fluent and low-latency while maintaining structured internal reasoning.
★ 171Forks 19
Towards Self-Evolving Proactive AI with Perpetual Memory
★ 218PythonForks 22
Real-time Breeze-TTS-2 streaming on Apple Silicon — RTF from ~4 down to ~1. mlx-audio 0.5.1's depth decoder recomputes each frame's existing acoustic codes instead of reusing its KV cache. Adds intra-frame KV reuse, one CPU read-back per frame, an optional compiled whole-frame path, and RTF/underflow stats.
★ 77PythonForks 12
★ 29PythonForks 0
I'm Xie Zhifei
★ 0Forks 0
Measuring Massive-Computational Math Reasoning with Code in LLMs
★ 0HTMLForks 0