MadeBy561/MadeBy561
Lifelong developer. Local Inference Lab contributor, maker of Doorplate.
Lifelong developer. Proud contributor to Local Inference Lab, a non-profit building open tooling for running LLMs on NVIDIA Blackwell. Creator of Doorplate.
Lifelong developer. Local Inference Lab contributor, maker of Doorplate.
LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
Run NVIDIA's GLM-5.2 eval axes (GPQA Diamond, SciCode, ...) against any OpenAI-compatible endpoint — measure what a REAP prune/quant costs.
A high-throughput and memory-efficient inference and serving engine for LLMs