alexeldeib/double-dipper
Weekly AI-written recap for the Double Dipper fantasy league
Weekly AI-written recap for the Double Dipper fantasy league
Production-Grade Container Scheduling and Management
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
FlashInfer: Kernel Library for LLM Serving
CoreWeave Terraform Provider
CLI for interacting with incident.io for batch automation
Opt-in command doctrine for coordinating Codex, Claude, and subagents as fireteams, squads, and platoons.
The open source coding agent.
An open-source database of AI models.
MLCommons Inference Endpoints repository
A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
Supercharge Your LLM with the Fastest KV Cache Layer
A Datacenter Scale Distributed Inference Serving Framework
Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
TensorStore S3 write throughput optimization benchmarks for JAX arrays
Minimal tensorstore S3 roundtrip test for Coreweave vhost-style object storage
A markdown-based knowledge tracking CLI for projects
A CLI to estimate inference memory requirements for Hugging Face models, written in Python.
Terminal fireworks show - ASCII art celebration in your terminal
Elite football data analytics dashboard for fans
Next Generation Agentic Proxy for AI Agents and MCP servers