valarLip/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
Self-contained GPU kernel profiler for ROCm. Zero roctracer/rocprofiler-sdk dependency. HSA-only interception, SQLite trace output, Perfetto compatible.
SGLang is a fast serving framework for large language models and vision language models.
AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease of use and versatility of software and hardware.
Config files for my GitHub profile.