simon-mo/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
cofounder of @Inferact, lead maintainer of @vllm-project
A high-throughput and memory-efficient inference and serving engine for LLMs
Tar file reading/writing for Rust
S3 Service Adapter
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Development repository for the Triton language and compiler
Fast and memory-efficient exact attention
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
FlashInfer: Kernel Library for LLM Serving
Container Express is a tool to accelerate Docker push and pull.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
A lightweight server clone of Azure Storage that simulates most of the commands supported by it with minimal dependencies
This repository is for active development of the *unofficial* Azure SDK for Rust. This repository is *not* supported by the Azure SDK team.