wxsIcey/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
Community maintained hardware plugin for vLLM on Ascend
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Dockerfiles for Ascend CANN