vLLM Core Maintainer | MTS at Red Hat
Repositories
tlrmchlsmth/dotfiles
tlrmchlsmth/gpu-roofline-explorer
Tensor Core Sizes from Volta to Blackwell: interactive SMEM rooflines and operand reuse
tlrmchlsmth/oubliette
Disposable vCluster sandboxes for agent-spawned LLM-serving cliques
tlrmchlsmth/vllm-synthid
Multi-tenant SynthID Text watermarking plugin for vLLM
tlrmchlsmth/crucible
Crucible: an autonomous goal-directed research loop engine
tlrmchlsmth/agentx-mvp
tlrmchlsmth/nightly-eval
tlrmchlsmth/nixl-agent-reclaim-repro
tlrmchlsmth/CoasterBench
An open source re-implementation of RollerCoaster Tycoon 2 🎢
tlrmchlsmth/coolS
A LaTeX package to use the Cool S as a symbol in math equations
tlrmchlsmth/hermes-agent
The agent that grows with you
tlrmchlsmth/j-llm-d
Justfile harness for llm-d
tlrmchlsmth/claudectx
kubectx for AI coding agents — switch paired Claude Code + Codex CLI contexts (settings, tokens, skills, MCP servers) and translate config between them
tlrmchlsmth/tv
tlrmchlsmth/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
tlrmchlsmth/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
tlrmchlsmth/vllm-skills
tlrmchlsmth/flashinfer
FlashInfer: Kernel Library for LLM Serving
tlrmchlsmth/llm-d
llm-d is a Kubernetes-native high-performance distributed LLM inference framework
tlrmchlsmth/llmd-routing-bench
Benchmarking tool for the llm-d routing sidecar (P/D disaggregation overhead)
tlrmchlsmth/DeepEP
DeepEP: an efficient expert-parallel communication library
tlrmchlsmth/vllm-dev-env
tlrmchlsmth/combine_traces
tlrmchlsmth/prefill-decode-experiments
tlrmchlsmth/llm-d-dev-img
tlrmchlsmth/my_pods
tlrmchlsmth/guidellm
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs