Gregory-Pereira/flashinfer
FlashInfer: Kernel Library for LLM Serving
Sr. Machine Learning Engineer @ Red Hat | Inference Engineering | Building llm-d: distributed inference for LLMs on Kubernetes
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine for LLMs
llm-d is a Kubernetes-native high-performance distributed LLM inference framework
Manages Unified Access to Generative AI Services built on Envoy Gateway
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Inference scheduler for llm-d
Beads - A memory upgrade for your coding agent
Latency prediction service for ML-model based scoring with llm-d-inference-scheduler
Disconnected readiness scanner for RHOAI component repos — scores repos for air-gapped/disconnected deployment compatibility
A github workflow to automate the creation of bugs in jira by reading the snyk scan reports.
The batch gateway is an llm-d implementation of the OpenAI batch inference API
A vLLM plugin built on the FlagOS unified multi-chip backend.
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
Fast and memory-efficient exact attention
Community maintained hardware plugin for vLLM on Ascend
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
repo for CI and infrastructure required to maintain llm-d org member repos
Collection of demos for building Llama Stack based apps on OpenShift
Central repository for managing Konflux resource files. This streamlines maintenance by consolidating configurations and leveraging GitHub Actions for automated syncing.
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Agent skills for vLLM
To centrally store all the Konflux artifacts
📜Fork for tracking CNCF projects
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
Variant optimization autoscaler for distributed inference workloads