adulbrich
## Problem `guides/generative-ai.mdx` (3,429 words) was last swept in #243, before #343 rewrote `ai-project-setup.mdx`. Its sourcing holds up (every cited claim sits inside its registry entry), but its structure has three problems. 1. **It carries the catalog #342 cut from ai-project-setup.** `## AI Coding Tools` and `## Choosing the Right Tool` (about 600 words, lines 106 to 153) are four product lists, a dozen vendor links, two benchmark trackers, and selection criteria (context window, speed) that date on a vendor's release schedule. The guides skill names the fix: cut whatever dates fastest, which is almost always a tool catalog. 2. **It describes ai-project-setup as it was.** - Line 86 says ai-project-setup shows the confidentiality rule "in both its team-readable and its tool-readable form". The rewrite shows only the tool form (the never tier) and sends the people form to `working-agreement#ai-tool-usage`. - Line 157, the first Best Practices bullet, tells the reader to configure the tool for the repository (ai-project-setup's subject, which this page's opener says it doesn't cover) and links the Skills activity as the walkthrough, which covers skills only. - `ai-project-setup.mdx:86` links `generative-ai#start-with-clear-intent` for "the task names its finish line". That section is about being specific and never mentions a finish line. 3. **It restates itself and lists what it should explain.** "Don't commit code you can't explain" has four homes (tl;dr, line 22, line 57, line 76). The three fragment lists under What AI Can and Cannot Do (lines 30 to 58) are the list-of-fragments pattern the skill bans; the paragraph above them already carries the usable test. Separately, the page gives a student with no budget one option (Copilot Student) and doesn't tell the reader how varied the models are or how to compare them. Many labs now publish capable models across a wide range of prices, several of them open-weight, and which one is the best value changes every few months. arena.ai publishes Pareto frontier charts ([code](https://arena.ai/leaderboard/code/pareto), [agent](https://arena.ai/leaderboard/agent/pareto), plus text, vision, search, document) that plot performance against price, so a reader can see the current frontier for themselves instead of trusting a page that names one. Open harnesses such as [OpenCode](https://opencode.ai/docs/) connect to any provider with an API key and read the same `AGENTS.md` ai-project-setup teaches, so a team's repository setup carries over whatever model each member picks. `guides/working-agreement.mdx`, AI Tool Usage (line 102 onward), still points at the generative-ai guide for shared configuration and names `CLAUDE.md` as the example file. ## Direction Organize generative-ai around its existing idea, which is sound: validate the hard-to-reverse decisions yourself and delegate the rest as far as your checks catch. Cut what dates fastest, and replace the catalog with a short section on choosing a model and a harness by comparing them, never by vendor. Teach the method, not the current answer. The page says that models vary widely in capability and price, that the best value moves, and how to compare them: benchmarks, Pareto curves, and a trial on your own repository. It names no model as expensive or cheap, better or worse, and states no price, plan feature, or ranking, because each of those is true today and wrong within months. The same rule applies to what the page already says: a present-tense claim about a vendor's plan or terms goes, and the principle behind it stays. Keep the automation table. `docs/decisions/2026-08-17-four-skills-assessment-design.md` §4.3 makes this page the "what can I hand off?" hub; refresh its rows rather than cut it. Bring in the practitioner and research voices the page is missing, each where it replaces a list or backs a habit (see Sources to add). None of them overlaps ai-project-setup: that page owns how the repository lets an agent check its own work; this page owns how you check, question, and learn from the tool. Target about 3,300 words. ## Outline 1. **Opening and tl;dr.** Unchanged in substance. Add one bullet: compare models on benchmarks and a price-performance curve, check who hosts your data, then try your shortlist on your own repository. 2. **What AI Can and Cannot Do.** Keep the wide-answer-space paragraph, the hallucination paragraph, and the package-name evidence (Spracklen). Cut the three fragment lists and put the jagged frontier (Dell'Acqua) in their place: the boundary between what a model does well and badly is uneven, moves with every release, and is found by testing your own tasks; outside it, AI made skilled people worse than working alone. Say the study was consulting work, not software. Fold the AI search point ("open each cited link") into the hallucination paragraph. Add sycophancy (Sharma) as its own short paragraph: models tilt toward what the user seems to believe, so ask neutral questions and ask for the case against your plan. 3. **Using AI Effectively.** - `### Start with clear intent`: keep the heading (inbound anchor from ai-project-setup). Add the finish line: say what done looks like, and the check that proves it, before you hand off the task. Link ai-project-setup's Define Done section for the repository side; don't repeat Osmani's example. - `### Verify what matters, and build the checks...`: keep, and add Willison's argument that you deliver code you've proven to work, including having seen it work yourself, because otherwise the verification lands on your reviewer. This is the author's proof to the team; ai-project-setup's "ask for the test output" is the agent's proof to the author. Don't merge the two. - Add vibe coding, in the verify section or its own short subsection: define it with Karpathy's coinage (accept every change without reading the diff), give Willison's argument that most AI-assisted programming isn't that and Beck's augmented coding as the disciplined alternative, and say where each fits: vibe coding for throwaway prototypes, augmented coding for anything others build on. - `### Use AI to learn`: keep as its own section (this reverses the earlier proposal to merge it) and ground it in Shen and Tamkin: full delegation, progressive reliance, and iterative AI debugging went with the lowest mastery scores; asking for explanations or asking only conceptual questions while writing the code yourself went with the highest. "Don't commit code you can't explain" lives here, once, plus the tl;dr; cut its other body copies. - `### Know when to stop prompting`: add starting a fresh session with what you learned, since a retry in the same context carries the wrong turns forward and output quality drops as the input grows (Chroma's Context Rot report, optional, named as an industry report). - `### Check what you paste`: keep the heading (inbound anchors). Fix line 86: the rule for people lives in the working agreement; the tool-readable form is `AGENTS.md`'s never tier in ai-project-setup. Cut the OpenAI plan specifics on line 84 (which plans train by default); keep the principle that terms differ by plan and you read the ones for yours. 4. **What You Can Automate.** Keep. Repoint the "Lint and format gates" row to `ai-project-setup#hooks-and-gates`. Say "checks" in the two table rows that say "gates". 5. **Choosing a Model and a Harness** (replaces AI Coding Tools and Choosing the Right Tool). - The autocomplete-versus-agent paragraph (line 130), kept as is. - Model and harness are separate choices. The harness is the agent program; the model is what it calls. Some harnesses are tied to one vendor's models, and open ones such as OpenCode let you choose. An open harness that reads `AGENTS.md` lets each teammate bring a different model to the same repository. Link ai-project-setup. - The variety: many labs publish capable models across a wide range of prices, some with open weights, and which is the best value for a given task changes often. No model named, no price, no ranking. - How to compare: benchmarks measure one kind of task each, so read what a benchmark tests before you read its ranking. A Pareto curve plots performance against price; the models on it are the ones nothing else beats on both at once. Link arena.ai's charts (and keep the benchmark trackers already on the page). Shortlist from the curve at your budget, then try the shortlist on a task from your own repository, because a leaderboard score comes from other people's prompts. Recommendation with its reason: prefer the cheapest model that passes your own checks, because the checks, not the model, decide how far you can delegate. - Where your data goes depends on who hosts the model, not who trained it. An open-weight model can run on the lab's own service, a third-party host, or your own machine, and each has different data terms. Some providers offer a lower price in exchange for training on your prompts, and some governments and institutions bar specific vendors' services on their devices, so read the terms and the rules that govern your work. Link back to Check what you paste. No vendor named. - The Copilot Student aside becomes one general sentence: several vendors offer free or discounted plans to verified students, so check before you pay. No plan named, no feature list. - Product names that don't appear above move to one line under Additional Readings. 6. **Best Practices.** Bullet 1 links ai-project-setup instead of the Skills activity. Keep "Prefer the cheapest tool that does the job". 7. **Some Truths.** Add Lee et al. next to Perry's paragraph: across 319 knowledge workers, more confidence in the AI went with less critical thinking and more self-confidence with more. Say it's a self-report survey. Optionally attribute the bottleneck truth to Karpathy's generate-then-verify loop (the way to go faster is to shorten verification). 8. **Industry and Academia, References, Additional Readings.** Unchanged, except: add arena.ai and OpenCode docs to Sources; add "Wire an Accessibility Audit Into CI" to Activities, since the automation table indexes accessibility auditing. ### `guides/working-agreement.mdx`, AI Tool Usage - First bullet: replace "(such as `CLAUDE.md`, agent skills, or a shared system prompt)" with one `AGENTS.md` checked into the repository, linked to `/guides/ai-project-setup/`. Add that the team agrees which model providers its confidentiality rule allows, since teammates may now bring different models to the same repository. - Replace the Generative AI LinkCard (line 122) with one to AI Project Setup, described as what to commit so every coding agent on the team works from the same rules. The team norms here are about shared configuration; individual habits are not the working agreement's concern, and generative-ai's opener already links back here. - Additional Readings: the "Set Up Your Repository's Skills" entry says it covers "the shared AI configuration and confidentiality rules the AI section points at"; the activity covers skills. Reword to what it does. ## Acceptance - [ ] No product list remains in the generative-ai body. No model is named as better, worse, cheaper, or more expensive than another; no price, plan feature, or ranking appears. The only product names in the body are harness examples, and the rest sit in Additional Readings. - [ ] Every present-tense claim about a vendor's plan, terms, or pricing is gone, including line 84's OpenAI specifics and the Copilot plan's feature list; the principle behind each stays. - [ ] Choosing a Model and a Harness says models vary widely and the best value moves, tells the reader to compare benchmarks and Pareto curves and then try the shortlist on their own repository, links arena.ai's charts and OpenCode, and states the hosting rule (host, not lab, decides where data goes) with a link to Check what you paste. - [ ] `#start-with-clear-intent` names a finish line, so `ai-project-setup.mdx:86` describes what its link lands on. - [ ] Line 86's confidentiality pointer sends each form of the rule to its owner. - [ ] "Don't commit code you can't explain" appears once in the body (Use AI to learn), plus the tl;dr. - [ ] `shen-2026` (preprint), `dellacqua-2026`, `sharma-2024`, and `lee-2025` registered in `src/data/sources/` and cited with `<Cite>`; each claim's locator read in full text before `verified: full-text`, otherwise `abstract` and the page says no more than the abstract. Confirm the ids against the first authors' family names when registering. - [ ] Karpathy, Willison, Beck, and (if used) Osmani are linked as arguments ("Willison argues"), never cited as Evidence. Chroma, if used, is registered as an industry report and named as one. - [ ] The jagged-frontier paragraph replaces the capability lists and says the study was consulting tasks; the Lee paragraph says self-report. - [ ] No new sentence overlaps ai-project-setup: no fresh-context reviewer, no agent self-checks, no `AGENTS.md`, no MCP. - [ ] Inbound anchors unchanged: `#start-with-clear-intent`, `#check-what-you-paste`. Any new anchor into ai-project-setup is confirmed by the build, not guessed. - [ ] working-agreement AI Tool Usage names `AGENTS.md`, links ai-project-setup in the bullet and the LinkCard, and no longer links generative-ai. - [ ] Guide skill's closing four sections in order; the three standalone greps return nothing that needs fixing. - [ ] `wc -w` on generative-ai about 3,300, and under 3,600. - [ ] `npm run build`, `canvas:export -- --strict`, `validate:sources`, `validate:activities`, `validate:dashes`, `validate:dates`, and `check:prose` pass. ## Open questions for review 1. Should the working agreement keep a second LinkCard to generative-ai, or does the single pointer to ai-project-setup cover it? This issue proposes replacing it. 2. Oregon added DeepSeek to its list of products barred on state IT. Does that reach OSU-owned devices or networks? If so, the course-level statement belongs in the syllabi or the introduction's AI policy, not the guide, which stays general ("read the rules that govern your work"). ## Sources to read - arena.ai, [Code Pareto](https://arena.ai/leaderboard/code/pareto) and [Agent Pareto](https://arena.ai/leaderboard/agent/pareto): what the axes and the price measure are, so the page describes how to read the chart correctly. Linked as a Reference, not cited as Evidence, and nothing it currently shows goes on the page. - OpenCode, [docs](https://opencode.ai/docs/): connects to providers by API key and reads `AGENTS.md`. ## Sources to add Evidence, registered and cited: - Shen and Tamkin (2026, preprint), [How AI Impacts Skill Formation](https://arxiv.org/html/2601.20245v1), with Anthropic's [summary](https://www.anthropic.com/research/AI-assistance-coding-skills). Randomized trial, 52 mostly junior developers learning the Trio library: the AI group scored 17% lower on a quiz about concepts used minutes earlier, with no significant speedup. Usage patterns sorted the scores (below 40% for full delegation, progressive reliance, iterative AI debugging; above 65% for explanations and conceptual questions). Goes in Use AI to learn. - Dell'Acqua et al. (2026), [Navigating the Jagged Technological Frontier](https://pubsonline.informs.org/doi/full/10.1287/orsc.2025.21838), *Organization Science*; the full text is open as an [HBS PDF](https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf). Preregistered experiment, 758 BCG consultants: inside the frontier, quality up more than 30%; on a task chosen to sit outside it, AI users 19 percentage points less likely to be correct. Goes in What AI Can and Cannot Do. - Sharma et al. (2024), [Towards Understanding Sycophancy in Language Models](https://proceedings.iclr.cc/paper_files/paper/2024/file/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf), ICLR. Five assistants consistently tilted answers toward the user's apparent views; humans and preference models sometimes preferred a convincing sycophantic answer over a correct one. Goes in What AI Can and Cannot Do. - Lee et al. (2025), [The Impact of Generative AI on Critical Thinking](https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/), CHI. Survey of 319 knowledge workers, 936 examples: confidence in GenAI associated with less critical thinking, self-confidence with more. Goes in Some Truths. Arguments, linked: - Karpathy's coinage of vibe coding, as quoted in Willison, [Not all AI-assisted programming is vibe coding](https://simonwillison.net/2025/Mar/19/vibe-coding/). - Kent Beck, [Augmented Coding: Beyond the Vibes](https://newsletter.kentbeck.com/p/augmented-coding-beyond-the-vibes). - Simon Willison, [Your job is to deliver code you have proven to work](https://simonwillison.net/2025/Dec/18/code-proven-to-work/). - Optional: Karpathy's Software 3.0 talk on the generate-then-verify loop ([notes](https://www.latent.space/p/s3); link the talk itself if it's posted); Chroma, [Context Rot](https://www.trychroma.com/research/context-rot) (industry report); Addy Osmani, [The 70% problem](https://addyo.substack.com/p/the-70-problem-hard-truths-about) (overlaps Shen and Tamkin, which is stronger evidence). Left out on purpose: Hashimoto, Böckeler, and Anthropic's Claude Code best practices (ai-project-setup's sources); GitClear's churn reports (disputed method); Mollick's *Co-Intelligence* rules (Dell'Acqua is the evidence under them).