Xarbirus/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of Facebook's LLaMA model in C/C++
Unsupervised text tokenizer focused on computational efficiency
Tensor library for machine learning