liz-badada/aisimulate
AISimulate predicts LLM serving behavior and searches for strong deployment configurations offline, without bringing up a GPU serving cluster
Hahahahahahahaha
AISimulate predicts LLM serving behavior and searches for strong deployment configurations offline, without bringing up a GPU serving cluster
SGLang is a fast serving framework for large language models and vision language models.
A set of examples based on verl for end-to-end RL training recipes.
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
FlashInfer: Kernel Library for LLM Serving
A Datacenter Scale Distributed Inference Serving Framework
A framework for efficient model inference with omni-modality models
<Beat AI> 又名 <零生万物> , 是一本专属于软件开发工程师的 AI 入门圣经,手把手带你上手写 AI。从神经网络到大模型,从高层设计到微观原理,从工程实现到算法,学完后,你会发现 AI 也并不是想象中那么高不可攀、无法战胜,Just beat it !
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Offline optimization of your disaggregated Dynamo graph
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.
A high-throughput and memory-efficient inference and serving engine for LLMs
AI agents running research on single-GPU nanochat training automatically
Benchmark SGLang on SLURM
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
Helpful kernel tutorials and examples for tile-based GPU programming
CUDA Templates and Python DSLs for High-Performance Linear Algebra
AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术
A Quirky Assortment of CuTe Kernels
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
DeepEP: an efficient expert-parallel communication library
Analyze computation-communication overlap in V3/R1.