SuperMarioYL/inference-cookbook
inference cookbook / inference 框架原理解析
inference cookbook / inference 框架原理解析
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
capforge is the just-in-time capability forge for agentic coding agents
Run-level exposure gauge for coding agents on local models — tracks wall-clock, state mutations, and token burn since last human review, forces a pause at thresholds.
One-command MCP-server installer for every coding-agent client (Claude Code, Cursor, Codex): detect, backup, idempotent write, handshake smoke test. v0.6.0 hardens input handling (catalog --args now errors, --clients ids validated) and runs one handshake per add instead of one per client.
agentq
Consumer AI that remembers your whole life — affordable on DeepSeek prefix-cache
A rehearsed undo ledger for developers running coding agents with write access
Hindsight: Agent Memory That Learns
面向 AI 编程开发者的指令文件预检红绿灯工具
A prefix-cache profit-and-loss layer for DeepSeek coding agents — wrap your client in two lines, see per-request cache HIT/PARTIAL/MISS, the ¥ each prefix-bust wasted, and the reorder that restores the discount.
CoverLock — lock an account-level cover style-pack (model, prompt scaffold, palette, layout) as a hash-locked YAML, then render size-compliant, safe-zone-aware Xiaohongshu cover sets with a self-proving consistency gallery. Runs fully offline with --model mock.
Reserve estimated per-task spend before wrapped LiteLLM calls, settle usage afterward, and block requests that exceed configured ceilings.
Run a coding agent on your in-network GPU box from one command - mTLS thin client, on-box ReAct loop or Aider plugin, allowlisted egress proxy with append-only audit, bash isolated in a loopback-only netns. Data never leaves the box; every outbound dial is logged.
AttestLoad — attest-before-load gate for coding agents: sign AI agent Skills and MCP servers into verifiable SBOM + provenance attestations, and refuse to load unattested code.
我的图床
面向工程团队的编码智能体复用闸门,每个新建文件和符号都必须先给出可核验的复用理由
Deduplicate keyed agent steps and HTTP retries with cached results, a local proxy and optional JSON-file persistence.
Normalize GLM-style tool-call responses for OpenAI-compatible Python clients, including sync, async and streaming adapters.
Context-eviction CLI for coding agents — traces how far each tool output propagated across turns, ranks what is silently eating your context budget, and replay-proves which drops are safe before you cut them (temp=0 ablation replay).
Scan project and dependency prose for agent-directed instruction patterns, with source locations, rule explanations and CI exit codes.
段绘(DuanHui)— 把整篇中文文章按段落语义配上同风格白底手绘插图的命令行工具。v0.6.0 修复中文输入保真:GBK/GB18030 文章直接可读,分段清理块引用与无空格列表标记,硬换行中文不再产生多余空格。
Mac 用户的 AI 模型缓存管家,一条命令盘点全盘模型权重并给出可安全回收的空间
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
TokenSched 给 Claude Code 的 token 预算装上了一个 CPU 调度器:它按子任务期望值预分配预算、预测超支,并在 5 小时窗口耗尽前自动把低价值工作降级到 Haiku 或抢占——把硬截断变成可调度的软退让。
riskshape
mcp-conform
LedgerMem 用留一法重放消融(leave-one-out replay ablation)为智能体记忆做检索归因:逐轮证明哪些记忆行真正改变了动作,把只写不读的黑盒存储变成一份可信、可裁剪的记忆影响力账本。
RefusalScope