GOavi101/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
I am a Software Engineer and have knowledge of Linux, Kubernetes, Docker, and Go.
A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM plugin for Spyre based on torch-spyre
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Spyre
System Level Intelligent Router for Mixture-of-Models
Central repository for managing Konflux resource files. This streamlines maintenance by consolidating configurations and leveraging GitHub Actions for automated syncing.
Model Registry provides a single pane of glass for ML model developers to index and manage models, versions, and ML artifacts metadata. It fills a gap between model experimentation and production activities. It provides a central interface for all stakeholders in the MLOps lifecycle to collaborate on ML models.
A CLI for building container images on Kubernetes!
Standardized Serverless ML Inference Platform on Kubernetes