Skills to probe into agent tools and suggestions to fix them.
The first skill, mcp-probe-and-prompt, is a Claude Code skill that measures an
MCP server the way a benchmark would, grades it against a gate built from shipped
failures, and writes the prompt that turns it into a task-specific server in Nexla
MCP Studio — plus the short, human-voice prompts that make an agent use it well.
It came out of three head-to-head benchmarks (Google Ads, HubSpot, Avoma) where the same defects kept deciding the result: silent row caps reported as complete, pagination parameters that do nothing, and payloads the agent never sees. The skill turns that experience into a loop anyone on the team can run in a few minutes.
Probe → Gate → Prompt. Measure what the agent actually receives, check it against fourteen yes/no questions, then write a prompt with numbers in it.
git clone https://github.com/Debanitrkl/agent-fix-skill.git
cp -r agent-fix-skill/mcp-probe-and-prompt ~/.claude/skills/
bash ~/.claude/skills/mcp-probe-and-prompt/scripts/setup.shsetup.sh creates a small Python virtualenv with the mcp SDK (pinned <2). To
share the skill with everyone working in one repo, copy mcp-probe-and-prompt/ into
that repo's .claude/skills/ instead of ~/.claude/skills/.
Requirements: Claude Code, Python 3.10+, and Node/npx if the server you are probing
only speaks OAuth (the probe goes through mcp-remote for those).
Export the key the server needs, then paste an mcpServers config into Claude Code and
say any of these:
- "check this MCP server is working"
- "why is this server slow / expensive"
- "compare Nexla's server with the vendor's"
- "write a prompt for MCP Studio to make this task-specific"
- "give me three prompts my sales-ops lead can use that return consistent numbers"
Claude runs the prober, reads the report, grades it against the gate, and drafts the prompts. You can also run the prober directly:
export NEXLA_SERVICE_KEY=...
P=~/.claude/skills/mcp-probe-and-prompt
$P/.venv/bin/python $P/scripts/probe.py \
--url https://api-genai.nexla.io/mcp/service_key/<server-slug> \
--header "Authorization: Bearer $NEXLA_SERVICE_KEY" --out ./probe-<slug>It only calls tools whose names look read-only (read, list, get, search…) and
skips anything named create, update, delete, manage. It writes report.md and
probe.json with, per tool: latency, bytes the model actually sees, row count, whether
truncated tells the truth, and whether limit / offset / cursor change the
result (it hashes the returned IDs to find out). Three real reports are in
examples/.
Both prompts are one short message written the way you would ask a colleague. Each does three jobs; you check for them after writing, rather than filling in a template. Every Studio prompt in this project that produced a working build was a paragraph like the second one below.
To the agent that uses the server — which things count and which dates, the facts you want with units, and what to do when a result is cut short:
How many campaigns had at least one impression between 13 July and 11 August 2026? Give me the count and the campaign names. Use the exact numbers from the tools — if anything comes back truncated or a tool fails, tell me instead of estimating.
"Campaigns" alone returned 47; "campaigns with at least one impression" returned 6, which was the answer everyone meant. The last sentence is the one people drop: without it, an agent handed a truncated page reported the partial count as the total.
To MCP Studio — what the server is for and what is broken, what to build instead, and the number that proves it is fixed:
I want to rebuild our HubSpot win/loss server. Right now it hands back 200 deals and says that's everything, when we actually have 1,010 closed deals — so every number built on it is wrong. Rather than dumping deals for the agent to add up, give me one small tool per question we actually ask — win rate, deal size for won vs lost, loss reasons, owner performance — with the maths done on your side so each answer is a few hundred bytes, not half a megabyte. If a result is ever partial, say so and include the total count and a way to fetch the rest. Build it and check it against HubSpot yourself before handing it back — no need to stop and ask me in between — and paste the live numbers next to each check: all time should come to 1,010 closed (559 won, 451 lost), the last 90 days 206 (141 won, 65 lost), and the top loss reasons Competitor 18 and Price 12.
Not this, even though every fact in it is the same — it reads like a ticket, so it gets worked like one, and nothing in it tells the builder what the server is for:
Rebuild the HubSpot win/loss server: one tool per question, each a single fixed query aggregated server-side, under 4 KB per response. Never return a partial result with
truncated: false— returntotal_availableand a working cursor. Before publishing, paste live results for: full range → 1,010 closed (559 / 451); last 90 days → 206 (141 / 65).
Studio's builder runs on GPT-5.6, and OpenAI's own prompting guidance for that model asks
for exactly this shape: goal, context, constraints, required evidence, success criteria and
output format, each stated once, autonomy defined once, and no "think harder". The skill's
studio-prompts.md maps the paragraph
onto those ingredients.
If the paragraph has no number in it, it is not finished. Longer worked examples (a fresh
build, a follow-up, and a dynamic-pagination fix) are in
mcp-probe-and-prompt/references/examples/.
The prompt book docs/ask-then-branch.html (also in the skill as
references/studio-conversation.md) is for the moment you are sitting next to MCP Studio
watching it think. Studio shows three things while it works — a collapsible Thought for Ns
block, an N tools called group, and a one-line claim — and every prompt in the book answers
one of them. Each situation is a tree: what Studio said, three prompts you can send back, the
condition that makes the answer conclusive, and a root prompt that compares against a number
Studio did not produce.
The rule that runs through it: a list of maybes is a branch, not an answer. When Studio says "probably wrong id, different org, or the OAuth flow hasn't finished", the reply is:
Don't give me three maybes. Test each one: does 29226 exist in any org you can see, is it shared with this account, is there an unfinished OAuth flow? Tell me which one it is and what you ran to find out.
The trunks are quoted from real Studio sessions (account details elided); the "typical" ones are paraphrased patterns from the HubSpot builds in the benchmark.
| Path | What |
|---|---|
mcp-probe-and-prompt/SKILL.md |
the method Claude follows |
mcp-probe-and-prompt/scripts/probe.py |
the prober — streamable HTTP or stdio, read-only tools only, redacts credentials from its own output |
mcp-probe-and-prompt/references/gate.md |
the 14-line gate; every line is a defect a shipped build had |
mcp-probe-and-prompt/references/studio-prompts.md |
how to write the MCP Studio prompt: the paragraph a person sends, the GPT-5.6 mapping, do-not list, weak→strong words |
mcp-probe-and-prompt/references/agent-prompts.md |
the one-message agent prompt and its three jobs, examples that worked and failed, with results |
mcp-probe-and-prompt/references/case-studies.md |
Google Ads, HubSpot, Avoma numbers to quote |
mcp-probe-and-prompt/evals/evals.json |
test prompts for iterating on the skill |
examples/ |
real probe reports from three Nexla servers (HubSpot, Google Ads, Avoma) |
docs/probe-gate-prompt.html / .pdf |
the two-sheet brief for a mixed technical / non-technical audience |
docs/ask-then-branch.html |
the prompt book: what to say to Studio while it builds, as trees |
mcp-probe-and-prompt/references/studio-conversation.md |
the same prompts as text, with the ladder and the words that mean a branch isn't finished |
| Benchmark | Result |
|---|---|
| Google Ads — Nexla vs Google's official server | 1.00 vs 0.975 correctness; 1.0 vs 4.7 tool calls per task; 10.7k vs 50.1k tokens per task |
| HubSpot — Nexla vs HubSpot's hosted CRM MCP | 21.7k vs 709k tokens per task (32.7×), exact on both; four earlier Nexla builds failed on silent caps |
| Avoma — Nexla vs Avoma's official server | neither can answer the other's half: Avoma 1.00 / 0.80 / 0.00 and Nexla 0.33 / 0.00 / 0.50 across taxonomy / meetings / calls |
Details, including what went wrong in each run, are in
case-studies.md.
Keys are read from the environment (NEXLA_SERVICE_KEY or whatever the server needs).
Nothing in this repo contains a credential, and the prober redacts bearer tokens from
the command lines it records in its reports. Never paste a key into a report, a prompt,
or a chat — reference it as ${NEXLA_SERVICE_KEY}.