FisherKKK/slime
slime is an LLM post-training framework for RL Scaling.
slime is an LLM post-training framework for RL Scaling.
个人技术知识库
FlashInfer: Kernel Library for LLM Serving
SGLang is a fast serving framework for large language models and vision language models.
Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding (https://arxiv.org/abs/2602.22647)
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Nano vLLM
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
DeepEP: an efficient expert-parallel communication library
DeepGEMM: clean and efficient BLAS kernel library on GPU
learning by doing. Implement SOA kernels and RL algo to fully grasp performance limits
cuVS - a library for vector search and clustering on the GPU
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Transformer related optimization, including BERT, GPT
Efficient Triton Kernels for LLM Training
🚀 Efficient implementations for emerging model architectures
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
An Open Source Machine Learning Framework for Everyone
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Material for gpu-mode lectures
A high-throughput and memory-efficient inference and serving engine for LLMs
Fast and memory-efficient exact attention
CUDA Templates and Python DSLs for High-Performance Linear Algebra
knowledge base
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
A modern replacement for Redis and Memcached
An open-source C++ library developed and used at Facebook.