jdebache/flashinfer
FlashInfer: Kernel Library for LLM Serving
FlashInfer: Kernel Library for LLM Serving
nvidia-modelopt is a unified library of state-of-the-art model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM or TensorRT to optimize inference speed.
DeepEP: an efficient expert-parallel communication library
NVIDIA Inference Xfer Library (NIXL)
A Datacenter Scale Distributed Inference Serving Framework
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
A high-throughput and memory-efficient inference and serving engine for LLMs
Common recipes to run vLLM
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A framework for few-shot evaluation of language models.
NumPy aware dynamic Python compiler using LLVM
A lightweight LLVM python binding for writing JIT compilers
FlashMLA: Efficient MLA kernels
CUDA Templates for Linear Algebra Subroutines
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Build system, successor to Buck
Fast C++ logging library.
MSCCL++: A GPU-driven communication stack for scalable AI applications
common in-memory tensor structure
CUDA Kernel Benchmarking Library
Starlark implementation of bazel rules for CUDA.
Goal: Enable awesome tooling for Bazel users of the C language family.
Bazel support for Visual Studio Code
C++ Rules for Bazel
An Open Source Machine Learning Framework for Everyone
NVIDIA® TensorRT™, an SDK for high-performance deep learning inference, includes a deep learning inference optimizer and runtime that delivers low latency and high throughput for inference applications.