Omer Aplatony

@omerap12 · User

GitHub profile ↗ · Compare

@kubernetes SIG Autoscaling TL

Tel Aviv, Israel44 followers20 repositories

Repositories

omerap12/modelexpress

Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.

★ 0PythonForks 0

omerap12/speculators

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

★ 0Forks 0

omerap12/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

★ 0ShellForks 0

omerap12/foundation

☁️♮🏛 This repo contains several documents related to the operation of the CNCF. File non-technical issues related to CNCF here.

★ 0Forks 0

omerap12/org

Meta configuration for Kubernetes Github Org

★ 0GoForks 0

omerap12/k8s.io

Code and configuration to manage Kubernetes project infrastructure, including various *.k8s.io sites

★ 0HCLForks 0

omerap12/aibrix

Cost-efficient and pluggable Infrastructure components for GenAI inference

★ 0GoForks 0

omerap12/scaling-book

Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs

★ 0Forks 0

omerap12/karpenter

Karpenter is a Kubernetes Node Autoscaler built for flexibility, performance, and simplicity.

★ 0GoForks 1