University of Science and Technology Beijing
Repositories
Zth9730/awesome-ai-research-writing
Elevate your AI research writing, no more tedious polishing ✨
Zth9730/claw-code
The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools that make real things done. Now writing in Rust using oh-my-codex.
Zth9730/LongCat-Audio-Semantic
Zth9730/latent-diffusion
High-Resolution Image Synthesis with Latent Diffusion Models
Zth9730/RAE
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
Zth9730/transformer-vocos
Zth9730/usm-tokenizer
semantic tokenizer for speech and music
Zth9730/sequence-vector-quantize
dh vq-q or vae exp
Zth9730/audio-pipeline
Zth9730/blsp
BLSP: Bootstrapping Langauge-Speech Pre-training via Behavior Alignment of Continuation Writing
Zth9730/RepCodec
Models and code for RepCodec: A Speech Representation Codec for Speech Tokenization
Zth9730/SpeechTokenizer
This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples are presented on
Zth9730/wenet
Production First and Production Ready End-to-End Speech Recognition Toolkit
Zth9730/NLP-Models-Tensorflow
Gathers machine learning and Tensorflow deep learning models for NLP problems
Zth9730/NeMo-text-processing
NeMo text processing for ASR and TTS
Zth9730/awesome-source-free-test-time-adaptation
A curated list of papers in Test-time Adaptation, Test-time Training and Source-free Domain Adaptation
Zth9730/MS-SNSD
The Microsoft Scalable Noisy Speech Dataset (MS-SNSD) is a noisy speech dataset that can scale to arbitrary sizes depending on the number of speakers, noise types, and Speech to Noise Ratio (SNR) levels desired.
Zth9730/icefall
Zth9730/MaTe3D
MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing
Zth9730/faster-whisper
Faster Whisper transcription with CTranslate2
Zth9730/Unconstrained-AVSR
Zth9730/WavAugment
A library for speech data augmentation in time-domain
Zth9730/PromptingWhisper
Promting Whisper for Audio-Visual Speech Recognition, Code-Switched Speech Recognition, and Zero-Shot Speech Translation
Zth9730/fairseq2
FAIR Sequence Modeling Toolkit
Zth9730/chirp
Zth9730/MyArxiv
Zth9730/RetNet
An implementation of "Retentive Network: A Successor to Transformer for Large Language Models"
Zth9730/Whisper-Finetune
微调Whisper语音识别模型和加速推理,支持Web部署和Android部署
Zth9730/unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities