Xie Zhifei

@xzf-thu · User

GitHub profile ↗ · Compare

Phd. Realtime-interaction. xzf-thu.github.io

227 followers11 repositories

Repositories

xzf-thu/VoiceMem

Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing information and emotions respectively, while a fully streaming architecture eliminates latency at the fundamental level.

★ 2,283PythonForks 173

xzf-thu/Mega-ASR

First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**

★ 1,146PythonForks 75

xzf-thu/Audio-Reasoner

The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.

★ 300PythonForks 24

xzf-thu/Mini-Omni-Reasoner

Mini-Omni-Reasoner: a real-time speech reasoning framework that interleaves silent reasoning tokens with spoken response tokens (“thinking-in-speaking”), exploiting the LLM–audio throughput gap to keep speech fluent and low-latency while maintaining structured internal reasoning.

★ 171Forks 19

xzf-thu/Pask

Towards Self-Evolving Proactive AI with Perpetual Memory

★ 218PythonForks 22

xzf-thu/BreezeTTS2_Mac_Streaming

Real-time Breeze-TTS-2 streaming on Apple Silicon — RTF from ~4 down to ~1. mlx-audio 0.5.1's depth decoder recomputes each frame's existing acoustic codes instead of reusing its KV cache. Adds intra-frame KV reuse, one CPU read-back per frame, an optional compiled whole-frame path, and RTF/underflow stats.

★ 77PythonForks 12

xzf-thu/MMRC

Measuring Massive-Computational Math Reasoning with Code in LLMs

★ 0HTMLForks 0