CloudRipple/sglang-omni
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
An Engineering Student @OpenMOSS , @sii-research and @ Fudan University
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
SGLang is a high-performance serving framework for large language models and multimodal models.
ComfyUI custom nodes for OpenMOSS MOSS-TTS v1.5 (Local-Transformer 48kHz stereo + Delay 8B 24kHz)
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
MOSS-TTSD is a spoken dialogue generation model that enables expressive dialogue speech synthesis in both Chinese and English, supporting zero-shot multi-speaker voice cloning, and long-form speech generation.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.