Repositories
Danielkinz/paged_exllamav2
A fast inference library for running LLMs locally on modern consumer-class GPUs
Danielkinz/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A fast inference library for running LLMs locally on modern consumer-class GPUs
A high-throughput and memory-efficient inference and serving engine for LLMs