omerap12/kubernetes
Production-Grade Container Scheduling and Management
@kubernetes SIG Autoscaling TL
Production-Grade Container Scheduling and Management
Enhancements tracking repo for Kubernetes
Autoscaling components for Kubernetes
Kubernetes-native Job Queueing
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
Test infrastructure for the Kubernetes project.
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Achieve state of the art inference performance with modern accelerators on Kubernetes
llm-d Router: The intelligent entry point for inference requests
☁️♮🏛 This repo contains several documents related to the operation of the CNCF. File non-technical issues related to CNCF here.
A Python script for downloading comic books from readcomicsonline.ru.
Meta configuration for Kubernetes Github Org
Kubernetes community content
Config files for my GitHub profile.
Code and configuration to manage Kubernetes project infrastructure, including various *.k8s.io sites
Cost-efficient and pluggable Infrastructure components for GenAI inference
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
Kubernetes website and documentation repo:
Karpenter is a Kubernetes Node Autoscaler built for flexibility, performance, and simplicity.