ispobock/InferenceX
Open Source Continuous Inference Benchmarking Qwen3.5, DeepSeek, GPTOSS - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
Building @sgl-project
Open Source Continuous Inference Benchmarking Qwen3.5, DeepSeek, GPTOSS - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Fast and memory-efficient exact attention
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3.
Kernel Library Wheel for SGLang
Materials for learning SGLang
Structured Text Generation
FlashInfer: Kernel Library for LLM Serving
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
SGLang is a structured generation language designed for large language models (LLMs). It makes your interaction with models faster and more controllable.
A high-throughput and memory-efficient inference and serving engine for LLMs
OpenCompass is an LLM evaluation platform, supporting a wide range of models (InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Curated list of project-based tutorials
PyTorch implementation for image classification on MNIST/CIFAR10/trashnet.
C++ implementation for the Harris Corner Detection algorithm