Artifacts from Tim Nunamaker's AI coding workshop (2026-04-04). Built live during workshop prep in about two hours, using the exact loop the workshop teaches.
A Claude Code skill that acts as Tim's proxy during the workshop. When an attendee's coding agent loads this skill, it guides them through picking a project, prompting effectively, and getting unstuck — in Tim's voice, grounded in real patterns mined from his own Claude session history.
The skill encodes:
- The loop: interview → prior art → plan to disk → validation criteria up front → build → verify empirically
- Tim's code-quality lens: essential vs incidental complexity (Hickey), one-caller helpers get inlined (Carmack), lean functional (Clojure-influenced)
- Unstick patterns for common situations: circling agents, AI slop, ambition without validation, "I don't know what to prompt"
- Experience-level calibration across beginner, intermediate, and advanced attendees
- Verbatim Tim quotes from his real prompt corpus
Install:
cp -r skill/workshop-guide ~/.claude/skills/A separate Claude Code skill for conducting rigorous, source-verified market research and producing executive-shareable artifacts (CSV, landscape report, executive memo). Governing principle: be a mirror, not an advisor. Includes supporting Python scripts for scaffolding, URL verification, findings merging, and linting.
Install:
cp -r skill/market-research ~/.claude/skills/A tool that extracts user prompts from Claude Code session logs (~/.claude/projects/), stores them in SQLite with FTS5, and runs non-LLM analysis (frequency, patterns, recency-weighted samples) to surface how you actually prompt. Built live to ground the workshop-guide skill in Tim's real voice rather than his self-report of it.
cd prompt-mine
python3 extract.py # parse ~/.claude/projects/**/*.jsonl → prompts.db
python3 analyze.py # generate insights.md (frequency, patterns, samples)
python3 curate.py # generate curated.md (recency-weighted, noise-filtered)Scale: extracted 9,029 human prompts from 1,580 session files (~2.5GB) in under a minute. All non-LLM analysis — frequency, regex patterns, recency tiering. No embeddings, no pgvector, no LLM calls over the full corpus.
Privacy: the extracted data (prompts.db, curated.jsonl, curated.md, insights.md, highlights.jsonl) is gitignored. Run it on your own logs.
Empirical evidence that the workshop-guide skill outperforms generic Claude. 5 realistic workshop scenarios, each run twice (with skill loaded vs baseline), graded against 5-7 assertions per scenario:
| Eval | With Skill | Baseline | Word count (skill / baseline) |
|---|---|---|---|
| Blank-slate beginner picking a project | 5/5 | 2/5 | 185 / 628 |
| Agent going in circles on a bug | 6/6 | 4/6 | 196 / 884 |
| Picked project, now what? | 7/7 | 0/7 | 277 / 694 |
| Output feels like AI slop | 5/5 | 2/5 | 234 / 569 |
| Ambitious advanced attendee | 6/6 | 2/6 | 622 / 1,299 |
| Total | 29/29 (100%) | 10/29 (34%) | — |
Skill responses are 52–78% shorter than baselines on every eval and also run faster (26.3s vs 45.4s average). The largest gap is on "picked project, now what?" — baseline scores 0/7 because Tim's specific loop (interview → prior art → plan-to-disk → validation criteria) is a Tim-specific pattern that generic Claude does not suggest.
Each eval directory contains the full response, grading JSON, and timing data for both configurations.
This repo was built during workshop prep using the exact loop the workshop-guide skill teaches:
- Interview — long ChatGPT conversation about format, audience, and goals
- Prior art — studied how Simon Willison runs workshops, borrowed the budget-capped API key pattern
- Mine real data — built
prompt-mineto extract and analyze ~9k real prompts instead of relying on self-report - Plan to disk — detailed plan for what the skill needed to do, grounded in the mined patterns
- Validation criteria up front — 5 eval scenarios with 5–7 assertions each, defined before the skill was written
- Build —
workshop-guideskill with verbatim Tim quotes from the mined corpus - Verify empirically — 10 parallel subagent runs (with-skill and baseline), programmatic grading, benchmark
The point: if the skill can be built in two hours using its own advice, the advice works.
workshop/
├── skill/
│ ├── workshop-guide/ # workshop proxy skill (SKILL.md + eval prompts)
│ └── market-research/ # market research skill + supporting scripts
├── prompt-mine/ # corpus mining tool (data gitignored)
│ ├── extract.py # streaming JSONL → SQLite with FTS5
│ ├── analyze.py # frequency, patterns, samples
│ ├── curate.py # recency-weighted, noise-filtered
│ ├── grade.py # assertion grader for eval responses
│ └── make_benchmark.py # benchmark.json builder
├── evals/iteration-1/ # 5 scenarios × 2 configurations, graded
│ ├── benchmark.json
│ └── eval-*/
│ ├── with_skill/
│ └── without_skill/
└── README.md