lishunyang12/vllm-omni-pr-dashboard
vLLM-Omni PR dashboard: contributors, recent submissions, and topics
vLLM-Omni PR dashboard: contributors, recent submissions, and topics
Reproducible SM120 FlashInfer skip-softmax validation for vLLM-Omni: H3 and Wan videos, gated threshold 0.5 / until 0.94, Nsight traces and raw test data.
H3/FastH3 inference optimizations: AdaLN, MXFP8/SwiGLU, Sage/VSA, VAE and RDMA, with DGX Spark experiment guidance.
FastH3 single-factor lossy ablations: seven prompts, five variants, 35 original videos.
OpenVDN single-factor lossy ablations: 40 videos, 2 baseline repeats, and 32 paired comparisons.
A unified inference and post-training framework for accelerated video generation.
Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.
vLLM Quantization plugin for GGUF
Model-agnostic SVDQuant toolkit for diffusion transformers
lishunyang 的技术博客 — vLLM-Omni / 推理
Common recipes to run vLLM
A Datacenter Scale Distributed Inference Serving Framework
Fast and memory-efficient exact attention
Web playground (vLLM-Omni plugin) for NVIDIA Cosmos3 world-foundation models
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
This Python-based tool provides visualization support for Cadence Virtuoso analog simulations by generating Voltage Transfer Characteristic (VTC) graphs and calculating Static Noise Margin (SNM) for SRAM cells.
A framework for efficient model inference with omni-modality models
how to optimize some algorithm in cuda.
A high-throughput and memory-efficient inference and serving engine for LLMs
FlashInfer: Kernel Library for LLM Serving
Flash-Attention-3 forward kernel
a collection of skills for vllm-omni
SOTA rounding-based quantization for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
SGLang is a high-performance serving framework for large language models and multimodal models.