Blocked by: #4
**Placeholder, deliberately not a spec.** Deferred from the 2026-09-15 grilling
session that produced #4 so it would not be forgotten. Needs its own design pass
before it is `ready-for-agent`.
## The gap
`agentdocs lookup` (#4) answers only *exact* lookups: the caller knows the gem
and the fully-qualified entity. A real fraction of lookups start fuzzier — "which
method does X", "what was that class called". `SKILL.md` serves those today with
two grep recipes:
```
grep -i <term> index.md
grep -rn '^### #method_name' .
```
Because `search` is deferred, those recipes stay in the skill, and nothing else
can own them. That is the main reason `SKILL.md` lands around 35 lines after #4
rather than the ~18 that moving all mechanics out would have bought.
## Hard constraint already settled
**Exact lookup must never silently fall back to fuzzy matching.** A guess that
reads as an answer is exactly the plausible-wrong-answer failure #4 exists to
eliminate. So this is a **separate subtool**, not a `--search` flag on `lookup`,
and its results are labelled as *candidates*, never returned as an answer. The
two share one support class in `lib/`.
## Open
Ranking, whether the gem argument stays required, whether a name-only search can
span all resolved dependencies, and what the output shape is when several
candidates tie.
## Design notes, 2026-09-15 — still not a spec
Measured against the local corpus of 116 generated bundles (~46 MB, ~7,470 entity
files, ~59,900 members). Recording what the data says, and a corrected framing.
No implementation is proposed; we do not know enough yet.
### Corrected motivation
The original text cites shrinking `SKILL.md` as the payoff. That is the wrong
motivation and should not drive the design. The payoff is **reducing the work an
agent does**: if a common case can be closed in one deterministic call where the
skill's recipe currently costs N, that is the win. Line count in `SKILL.md` is a
side effect, not a goal.
**The format is not in scope.** Everything needed is already in a bundle. Whatever
this becomes, it extracts information that is present; it does not ask the format
to carry more.
### What the corpus says
Cost of the brute-force fallbacks, which the issue treated as an unknown:
| | median | p90 | max |
|---|---|---|---|
| `index.md` | 2.0 KB | 18 KB | 169 KB (rubocop) |
| one class's `## Member Summary` | 379 B | 1.7 KB | 42 KB |
| whole bundle | 136 KB | 1.1 MB | 5.2 MB |
90 of 116 indexes are under 8 KB; only 9 exceed 20 KB. So "just read the whole
index" costs roughly 500 tokens for the median gem — cheaper than the tool round
trip that would replace it. Same for member summaries. The expensive tail is
specific and short: rubocop, prism, yard, rbs, bundler, language_server-protocol,
rss, parser.
**40.2% of index entries and 45.3% of member-summary lines carry no description
at all** — genuinely undocumented entities, and the rate is wildly gem-dependent
(prism 10/273 bare; bundler 248/332; debug 74/78; irb 89/113). Any matcher built
on descriptions is blind to ~40% of the corpus, unevenly. Name matching works on
100%. This should be the number the design is organized around.
**Class-level fuzzy lookup already works. Member-level is where the recipe
breaks.** Spot-checked with realistic queries: `grep -i flag toys/index.md` and
`grep -i markup yard/index.md` both land the answer on the first try. "read the
request body" against rack fails, because the answer is `Rack::Request#body` and
(a) `index.md` covers classes only, (b) `Rack::Request`'s own summary lines are
bare (`- #ip`, `- #params`), and (c) `#body` is reachable only through the
names-only **Inherited & Mixed-in Members** list, by deliberate format design. The
description exists, in `Rack/Request/Helpers.md`; the path to it is a traversal,
not a search.
### What this implies
1. The tool sketched in the original text — wrap a grep over descriptions — is the
weaker half of the problem and the half the plain recipe already handles.
2. The real gap is **member-level discovery across a gem**, and it is a traversal
problem (one hop through the inherited/mixed-in list) more than a matching one.
3. If a matcher is built, it should be **name-first, description-second**: exact
name, then name-fragment after stripping `#`/`.` sigils and splitting on `::`
and `_`, then description substring. The name tiers work on the whole corpus;
the description tier is a bonus on the documented 60%.
4. Ranking beyond those tiers is not worth building. The calling model is the
ranker; the job is to hand it a small, complete, uniformly shaped candidate
list — FQN, defining file, summary if one exists — in the shape `lookup` takes
as input. That also dissolves the "what if candidates tie" question below:
don't break ties, list them. They are candidates, not an answer.
### One separable win, split out as #10
The fuzzy recipe in `SKILL.md` opens by spending a whole `lookup` call on a
namespace the agent already knows, purely to read the bundle directory out of the
`Bundle:` header. A dedicated `agentdocs path GEM` would give that one bare,
shell-composable line instead, which matters mainly because direct grep is
first-class per `CLAUDE.md` and currently has no cheap entry point.
Filed separately as #10, since it needs none of this issue's open questions
answered. One dependency runs the other way: if `search` ships and covers the
common discovery cases, #10 stops serving *this* workflow and serves ad-hoc
exploration only.
### Still unknown
**Is the cross-gem case real?** "Which of my dependencies gives me X" is the one
thing an agent genuinely cannot do with grep — it does not know which bundles
exist, where they live, or which versions resolve. `DependencyResolver` does. If
that case is real it is the primary justification for a tool at all, and the
"whether the gem argument stays required" line below stops being a detail and
becomes the central question. It also carries a cost worth stating: spanning
resolved dependencies means every one of them must be *built*, which on a large
Gemfile is a lot of first-run minutes. Currently no evidence either way.
**How often does the fuzzy path get taken, and where does the recipe actually
fail?** Real usage telemetry would answer this, but there is no deployment, so
waiting on it blocks indefinitely. The cheap substitute is an **offline recall
eval** over the corpus already on disk: assemble N realistic "how do I…" questions
with known-correct FQN answers, then measure recall and token cost for each
strategy — grep the index, read the index whole, grep `### ` headings, grep the
whole bundle, and any proposed tool. That converts the central unknown into a
number using assets that already exist, and it should land before this issue is
`ready-for-agent`.
### Unchanged
The hard constraint stands: exact lookup never silently degrades to fuzzy
matching, this stays a separate subtool rather than a `--search` flag, and results
are labelled candidates. Still open: ranking, whether the gem argument stays
required, whether a name-only search can span resolved dependencies, and the
output shape when several candidates tie.