key4ng/recipes
Common recipes to run vLLM
Nebulizing..
Common recipes to run vLLM
A tiny Claude-Code-style CLI that chats with multiple AI models, runs bash with guardrails like a brave little engineer, and occasionally learns new tricks through plugins.
Because sometimes you want gh, but your company uses Bitbucket. lite-bb is a gh-style minimal CLI for pull request creation, diffs, reviews, and comments — designed for humans, scripts, and LLM agents.
A high-throughput and memory-efficient inference and serving engine for LLMs
Shepherd Model Gateway
Claude Code skill for Bitbucket CLI (lite-bb) — the gh equivalent for Bitbucket
For when your wife/girlfriend's company uses Datadog, write access to API keys is disabled, and someone configured 100+ stupid/noisy SMS alerts that mostly just need an ack.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Train a 64M-parameter LLM from scratch in just 2h!
Achieve state of the art inference performance with modern accelerators on Kubernetes
A developer journey through using and building with OME
TokenSpeed is a speed-of-light LLM inference engine.
OME is a Kubernetes operator for enterprise-grade management and serving of Large Language Models (LLMs)
A Rust reimplementation of genai-bench for benchmarking LLM serving systems at high concurrency with accurate timing and industry-standard metrics.
A minimal learning-focused implementation of SMG's gRPC router — proxies OpenAI-compatible requests to SGLang inference backends in ~1000 lines of Rust.
KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale
GitHub’s official command line tool
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
SGLang is a fast serving framework for large language models and vision language models.
A Datacenter Scale Distributed Inference Serving Framework
TensorZero is an open-source stack for industrial-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluation, and experimentation.
Nano vLLM
Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization