zhewenl/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
@Inferact | @vllm-project
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
NVIDIA Inference Xfer Library (NIXL)
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-performance and light-weight router for vLLM large scale deployment
Measuring frontier coding agents on original, long-horizon engineering tasks
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Benchmark SGLang on SLURM
An agentic skills framework & software development methodology that works.
Common recipes to run vLLM
A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
SGLang is a fast serving framework for large language models and vision language models.
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
This repository hosts code that supports the testing infrastructure for the PyTorch organization. For example, this repo hosts the logic to track disabled tests and slow tests, as well as our continuation integration jobs HUD/dashboard.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
A Datacenter Scale Distributed Inference Serving Framework