alexm-redhat/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
@vllm-project
A high-throughput and memory-efficient inference and serving engine for LLMs
OpenShell is the safe, private runtime for autonomous AI agents.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Scripts for automatic trtllm vs vllm benchmarks
Common recipes to run vLLM
Easy-to-use autoML interface to optimize deep neural networks for better inference performance and a smaller footprint.
Compiler for Neural Network hardware accelerators
Libraries and state-of-the-art automatic sparsification algorithms to simplify and accelerate performance
CPU inference engine that delivers unprecedented performance for sparse models
Neural network model repository for highly sparse models and optimization recipes