Repositories
jingyu-ml/JavisDiT
[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
jingyu-ml/Automodel
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
jingyu-ml/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
jingyu-ml/flashinfer
FlashInfer: Kernel Library for LLM Serving
jingyu-ml/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
jingyu-ml/sglang
SGLang is a fast serving framework for large language models and vision language models.
jingyu-ml/diffusers
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
jingyu-ml/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
jingyu-ml/smoothquant
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models