jeejeelee/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
@Inferact | @vllm-project
A high-throughput and memory-efficient inference and serving engine for LLMs
FlashInfer: Kernel Library for LLM Serving
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
Common recipes to run vLLM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
DeepEP: an efficient expert-parallel communication library
Accessible large language models via k-bit quantization for PyTorch.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
CUDA Templates for Linear Algebra Subroutines