Debanitrkl/agent-fix-skill

Skills to probe into agent tools and suggestions to fix them

★ 0Forks 0PythonGitHub ↗Compare

README

agent-fix-skill

Skills to probe into agent tools and suggestions to fix them.

The first skill, mcp-probe-and-prompt, is a Claude Code skill that measures an MCP server the way a benchmark would, grades it against a gate built from shipped failures, and writes the prompt that turns it into a task-specific server in Nexla MCP Studio — plus the short, human-voice prompts that make an agent use it well.

It came out of three head-to-head benchmarks (Google Ads, HubSpot, Avoma) where the same defects kept deciding the result: silent row caps reported as complete, pagination parameters that do nothing, and payloads the agent never sees. The skill turns that experience into a loop anyone on the team can run in a few minutes.

The method in one line

Probe → Gate → Prompt. Measure what the agent actually receives, check it against fourteen yes/no questions, then write a prompt with numbers in it.

Install

git clone https://github.com/Debanitrkl/agent-fix-skill.git
cp -r agent-fix-skill/mcp-probe-and-prompt ~/.claude/skills/
bash ~/.claude/skills/mcp-probe-and-prompt/scripts/setup.sh

setup.sh creates a small Python virtualenv with the mcp SDK (pinned <2). To share the skill with everyone working in one repo, copy mcp-probe-and-prompt/ into that repo's .claude/skills/ instead of ~/.claude/skills/.

Requirements: Claude Code, Python 3.10+, and Node/npx if the server you are probing only speaks OAuth (the probe goes through mcp-remote for those).

Use

Export the key the server needs, then paste an mcpServers config into Claude Code and say any of these:

  • "check this MCP server is working"
  • "why is this server slow / expensive"
  • "compare Nexla's server with the vendor's"
  • "write a prompt for MCP Studio to make this task-specific"
  • "give me three prompts my sales-ops lead can use that return consistent numbers"

Claude runs the prober, reads the report, grades it against the gate, and drafts the prompts. You can also run the prober directly:

export NEXLA_SERVICE_KEY=...
P=~/.claude/skills/mcp-probe-and-prompt
$P/.venv/bin/python $P/scripts/probe.py \
  --url https://api-genai.nexla.io/mcp/service_key/<server-slug> \
  --header "Authorization: Bearer $NEXLA_SERVICE_KEY" --out ./probe-<slug>

It only calls tools whose names look read-only (read, list, get, search…) and skips anything named create, update, delete, manage. It writes report.md and probe.json with, per tool: latency, bytes the model actually sees, row count, whether truncated tells the truth, and whether limit / offset / cursor change the result (it hashes the returned IDs to find out). Three real reports are in examples/.

What "a prompt with numbers in it" looks like

Both prompts are one short message written the way you would ask a colleague. Each does three jobs; you check for them after writing, rather than filling in a template. Every Studio prompt in this project that produced a working build was a paragraph like the second one below.

To the agent that uses the server — which things count and which dates, the facts you want with units, and what to do when a result is cut short:

How many campaigns had at least one impression between 13 July and 11 August 2026? Give me the count and the campaign names. Use the exact numbers from the tools — if anything comes back truncated or a tool fails, tell me instead of estimating.

"Campaigns" alone returned 47; "campaigns with at least one impression" returned 6, which was the answer everyone meant. The last sentence is the one people drop: without it, an agent handed a truncated page reported the partial count as the total.

To MCP Studio — what the server is for and what is broken, what to build instead, and the number that proves it is fixed:

I want to rebuild our HubSpot win/loss server. Right now it hands back 200 deals and says that's everything, when we actually have 1,010 closed deals — so every number built on it is wrong. Rather than dumping deals for the agent to add up, give me one small tool per question we actually ask — win rate, deal size for won vs lost, loss reasons, owner performance — with the maths done on your side so each answer is a few hundred bytes, not half a megabyte. If a result is ever partial, say so and include the total count and a way to fetch the rest. Build it and check it against HubSpot yourself before handing it back — no need to stop and ask me in between — and paste the live numbers next to each check: all time should come to 1,010 closed (559 won, 451 lost), the last 90 days 206 (141 won, 65 lost), and the top loss reasons Competitor 18 and Price 12.

Not this, even though every fact in it is the same — it reads like a ticket, so it gets worked like one, and nothing in it tells the builder what the server is for:

Rebuild the HubSpot win/loss server: one tool per question, each a single fixed query aggregated server-side, under 4 KB per response. Never return a partial result with truncated: false — return total_available and a working cursor. Before publishing, paste live results for: full range → 1,010 closed (559 / 451); last 90 days → 206 (141 / 65).

Studio's builder runs on GPT-5.6, and OpenAI's own prompting guidance for that model asks for exactly this shape: goal, context, constraints, required evidence, success criteria and output format, each stated once, autonomy defined once, and no "think harder". The skill's studio-prompts.md maps the paragraph onto those ingredients.

If the paragraph has no number in it, it is not finished. Longer worked examples (a fresh build, a follow-up, and a dynamic-pagination fix) are in mcp-probe-and-prompt/references/examples/.

Talking to Studio while it builds

The prompt book docs/ask-then-branch.html (also in the skill as references/studio-conversation.md) is for the moment you are sitting next to MCP Studio watching it think. Studio shows three things while it works — a collapsible Thought for Ns block, an N tools called group, and a one-line claim — and every prompt in the book answers one of them. Each situation is a tree: what Studio said, three prompts you can send back, the condition that makes the answer conclusive, and a root prompt that compares against a number Studio did not produce.

The rule that runs through it: a list of maybes is a branch, not an answer. When Studio says "probably wrong id, different org, or the OAuth flow hasn't finished", the reply is:

Don't give me three maybes. Test each one: does 29226 exist in any org you can see, is it shared with this account, is there an unfinished OAuth flow? Tell me which one it is and what you ran to find out.

The trunks are quoted from real Studio sessions (account details elided); the "typical" ones are paraphrased patterns from the HubSpot builds in the benchmark.

What's in the repo

Path What
mcp-probe-and-prompt/SKILL.md the method Claude follows
mcp-probe-and-prompt/scripts/probe.py the prober — streamable HTTP or stdio, read-only tools only, redacts credentials from its own output
mcp-probe-and-prompt/references/gate.md the 14-line gate; every line is a defect a shipped build had
mcp-probe-and-prompt/references/studio-prompts.md how to write the MCP Studio prompt: the paragraph a person sends, the GPT-5.6 mapping, do-not list, weak→strong words
mcp-probe-and-prompt/references/agent-prompts.md the one-message agent prompt and its three jobs, examples that worked and failed, with results
mcp-probe-and-prompt/references/case-studies.md Google Ads, HubSpot, Avoma numbers to quote
mcp-probe-and-prompt/evals/evals.json test prompts for iterating on the skill
examples/ real probe reports from three Nexla servers (HubSpot, Google Ads, Avoma)
docs/probe-gate-prompt.html / .pdf the two-sheet brief for a mixed technical / non-technical audience
docs/ask-then-branch.html the prompt book: what to say to Studio while it builds, as trees
mcp-probe-and-prompt/references/studio-conversation.md the same prompts as text, with the ladder and the words that mean a branch isn't finished

Where the numbers come from

Benchmark Result
Google Ads — Nexla vs Google's official server 1.00 vs 0.975 correctness; 1.0 vs 4.7 tool calls per task; 10.7k vs 50.1k tokens per task
HubSpot — Nexla vs HubSpot's hosted CRM MCP 21.7k vs 709k tokens per task (32.7×), exact on both; four earlier Nexla builds failed on silent caps
Avoma — Nexla vs Avoma's official server neither can answer the other's half: Avoma 1.00 / 0.80 / 0.00 and Nexla 0.33 / 0.00 / 0.50 across taxonomy / meetings / calls

Details, including what went wrong in each run, are in case-studies.md.

Credentials

Keys are read from the environment (NEXLA_SERVICE_KEY or whatever the server needs). Nothing in this repo contains a credential, and the prober redacts bearer tokens from the command lines it records in its reports. Never paste a key into a report, a prompt, or a chat — reference it as ${NEXLA_SERVICE_KEY}.

Contributors

Debanitrkl

Issues