SYLAR

@lishunyang12 · User

GitHub profile ↗ · Compare

National University of SingaporeShenzhen67 followers60 repositories

Repositories

lishunyang12/h3-optimizations

H3/FastH3 inference optimizations: AdaLN, MXFP8/SwiGLU, Sage/VSA, VAE and RDMA, with DGX Spark experiment guidance.

★ 1PythonForks 0

lishunyang12/FastVideo

A unified inference and post-training framework for accelerated video generation.

★ 0Forks 0

lishunyang12/fast-ulysses

Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.

★ 0Forks 0

lishunyang12/Lance

A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

★ 0Forks 0

lishunyang12/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 2PythonForks 1

lishunyang12/auto-round

SOTA rounding-based quantization for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

★ 0Forks 0

lishunyang12/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 0Forks 0

lishunyang12/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0Forks 0