googs1025/kvcached
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Cloud Native Beginner | AI-infra learninger | ex-ByteDance | @kubernetes Member | @aibrix maintainer | Kubernetes Contributor Award 2025 (Sigs-Scheduling)
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
A local-first routing and coordination engine for long-running agent work.
学摩尔线程 MUSA SDK 的学习记录
Kubernetes operator for deploying and managing OpenClaw AI agent instances with production-grade security, observability, and lifecycle management.
FlashInfer: Kernel Library for LLM Serving
GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
基于《代码整洁之道》《架构整洁之道》《程序员修炼之道》三本经典的中文代码审查 Claude Code Skill
CUDA Library Samples
Delivers efficient, stable, and secure data distribution and acceleration powered by P2P technology, with an optional content‑addressable filesystem that accelerates OCI container launch.
Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping
Nephio is a Kubernetes-based automation platform for deploying and managing highly distributed, interconnected workloads such as 5G Network Functions, and the underlying infrastructure on which those workloads depend.
LLM inference in C/C++
ANOLISA (Agentic Nexus Operating Layer & Interface System Architecture)
Envoy AI Gateway is an open source project for using Envoy Gateway to handle request traffic from application clients to Generative AI services.
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
llm-d Router: The intelligent entry point for inference requests
Achieve state of the art inference performance with modern accelerators on Kubernetes
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Variant optimization autoscaler for distributed inference workloads
A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Stateful API logic for agentic applications using vLLM
AVE 是面向安全运营的多源漏洞知识库,统一 AVE 编号并输出结构化 TOML,同时整理和校验公开 PoC/EXP 资产,支持按严重等级快速筛选高价值漏洞。漏洞爬取与梳理的代码和逻辑暂未开源。
vLLM 与 SGLang: 大模型高效推理双引擎实战配套代码
护网漏洞情报库 · 0day/1day/nday 漏洞数据 · 原腾讯文档迁移 · Issue 提交 · TOML 源数据