Make the world faster with GPUs
NVIDIA5 followers7 repositories
Repositories
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across SGLang, vLLM, TRT-LLM, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
★ 0RustForks 0
TokenSpeed is a speed-of-light LLM inference engine.
★ 0PythonForks 0
Rust multimodal preprocessing library from SMG
★ 0Forks 0
A Datacenter Scale Distributed Inference Serving Framework
★ 0Forks 0
BAC Demo KR
★ 2JavaScriptForks 1
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
★ 0PythonForks 0
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
★ 0Forks 0