Add an `agentdocs search` tool for discovery lookups

#5 · open · 1 comments

View on GitHub ↗

dazuma

Blocked by: #4 **Placeholder, deliberately not a spec.** Deferred from the 2026-09-15 grilling session that produced #4 so it would not be forgotten. Needs its own design pass before it is `ready-for-agent`. ## The gap `agentdocs lookup` (#4) answers only *exact* lookups: the caller knows the gem and the fully-qualified entity. A real fraction of lookups start fuzzier — "which method does X", "what was that class called". `SKILL.md` serves those today with two grep recipes: ``` grep -i <term> index.md grep -rn '^### #method_name' . ``` Because `search` is deferred, those recipes stay in the skill, and nothing else can own them. That is the main reason `SKILL.md` lands around 35 lines after #4 rather than the ~18 that moving all mechanics out would have bought. ## Hard constraint already settled **Exact lookup must never silently fall back to fuzzy matching.** A guess that reads as an answer is exactly the plausible-wrong-answer failure #4 exists to eliminate. So this is a **separate subtool**, not a `--search` flag on `lookup`, and its results are labelled as *candidates*, never returned as an answer. The two share one support class in `lib/`. ## Open Ranking, whether the gem argument stays required, whether a name-only search can span all resolved dependencies, and what the output shape is when several candidates tie.

Comments

dazuma

## Design notes, 2026-09-15 — still not a spec Measured against the local corpus of 116 generated bundles (~46 MB, ~7,470 entity files, ~59,900 members). Recording what the data says, and a corrected framing. No implementation is proposed; we do not know enough yet. ### Corrected motivation The original text cites shrinking `SKILL.md` as the payoff. That is the wrong motivation and should not drive the design. The payoff is **reducing the work an agent does**: if a common case can be closed in one deterministic call where the skill's recipe currently costs N, that is the win. Line count in `SKILL.md` is a side effect, not a goal. **The format is not in scope.** Everything needed is already in a bundle. Whatever this becomes, it extracts information that is present; it does not ask the format to carry more. ### What the corpus says Cost of the brute-force fallbacks, which the issue treated as an unknown: | | median | p90 | max | |---|---|---|---| | `index.md` | 2.0 KB | 18 KB | 169 KB (rubocop) | | one class's `## Member Summary` | 379 B | 1.7 KB | 42 KB | | whole bundle | 136 KB | 1.1 MB | 5.2 MB | 90 of 116 indexes are under 8 KB; only 9 exceed 20 KB. So "just read the whole index" costs roughly 500 tokens for the median gem — cheaper than the tool round trip that would replace it. Same for member summaries. The expensive tail is specific and short: rubocop, prism, yard, rbs, bundler, language_server-protocol, rss, parser. **40.2% of index entries and 45.3% of member-summary lines carry no description at all** — genuinely undocumented entities, and the rate is wildly gem-dependent (prism 10/273 bare; bundler 248/332; debug 74/78; irb 89/113). Any matcher built on descriptions is blind to ~40% of the corpus, unevenly. Name matching works on 100%. This should be the number the design is organized around. **Class-level fuzzy lookup already works. Member-level is where the recipe breaks.** Spot-checked with realistic queries: `grep -i flag toys/index.md` and `grep -i markup yard/index.md` both land the answer on the first try. "read the request body" against rack fails, because the answer is `Rack::Request#body` and (a) `index.md` covers classes only, (b) `Rack::Request`'s own summary lines are bare (`- #ip`, `- #params`), and (c) `#body` is reachable only through the names-only **Inherited & Mixed-in Members** list, by deliberate format design. The description exists, in `Rack/Request/Helpers.md`; the path to it is a traversal, not a search. ### What this implies 1. The tool sketched in the original text — wrap a grep over descriptions — is the weaker half of the problem and the half the plain recipe already handles. 2. The real gap is **member-level discovery across a gem**, and it is a traversal problem (one hop through the inherited/mixed-in list) more than a matching one. 3. If a matcher is built, it should be **name-first, description-second**: exact name, then name-fragment after stripping `#`/`.` sigils and splitting on `::` and `_`, then description substring. The name tiers work on the whole corpus; the description tier is a bonus on the documented 60%. 4. Ranking beyond those tiers is not worth building. The calling model is the ranker; the job is to hand it a small, complete, uniformly shaped candidate list — FQN, defining file, summary if one exists — in the shape `lookup` takes as input. That also dissolves the "what if candidates tie" question below: don't break ties, list them. They are candidates, not an answer. ### One separable win, split out as #10 The fuzzy recipe in `SKILL.md` opens by spending a whole `lookup` call on a namespace the agent already knows, purely to read the bundle directory out of the `Bundle:` header. A dedicated `agentdocs path GEM` would give that one bare, shell-composable line instead, which matters mainly because direct grep is first-class per `CLAUDE.md` and currently has no cheap entry point. Filed separately as #10, since it needs none of this issue's open questions answered. One dependency runs the other way: if `search` ships and covers the common discovery cases, #10 stops serving *this* workflow and serves ad-hoc exploration only. ### Still unknown **Is the cross-gem case real?** "Which of my dependencies gives me X" is the one thing an agent genuinely cannot do with grep — it does not know which bundles exist, where they live, or which versions resolve. `DependencyResolver` does. If that case is real it is the primary justification for a tool at all, and the "whether the gem argument stays required" line below stops being a detail and becomes the central question. It also carries a cost worth stating: spanning resolved dependencies means every one of them must be *built*, which on a large Gemfile is a lot of first-run minutes. Currently no evidence either way. **How often does the fuzzy path get taken, and where does the recipe actually fail?** Real usage telemetry would answer this, but there is no deployment, so waiting on it blocks indefinitely. The cheap substitute is an **offline recall eval** over the corpus already on disk: assemble N realistic "how do I…" questions with known-correct FQN answers, then measure recall and token cost for each strategy — grep the index, read the index whole, grep `### ` headings, grep the whole bundle, and any proposed tool. That converts the central unknown into a number using assets that already exist, and it should land before this issue is `ready-for-agent`. ### Unchanged The hard constraint stands: exact lookup never silently degrades to fuzzy matching, this stays a separate subtool rather than a `--search` flag, and results are labelled candidates. Still open: ranking, whether the gem argument stays required, whether a name-only search can span resolved dependencies, and the output shape when several candidates tie.