Your coding agent starts every session from zero, and you never see what it got wrong.
This finds out. It asks your agent five questions about your repo, cold, writes a small docs/ layer from your real code, then asks the same five again and prints the score.
The kit is one layer of the method I teach in Prompt to Production. One full module is free.
โญ If you find this useful, please star the repo so more developers can find it.
Your agent can read your code. It can't read the decisions behind it. Where does a new file go here? Why this database and not the obvious one? None of that is in the repo, so the agent guesses with whatever was most common in its training data.
So you write those decisions down where the agent will read them. By hand that's an afternoon per repo. This does it from your real code instead.
Not a Spec Kit replacement. Spec Kit writes a spec, then builds the code from it. This reads the code you already have and writes the memory your agent is missing. Use both.
Paste this to the agent already working in your repo:
Clone https://github.com/NirDiamant/Agentic_Engineering into a temp folder, read its RUN.md, and follow it on this repository.
That line works in Claude Code, Codex, Cursor and any agent that can clone and read. In Claude Code it is also a plugin, so the run is a slash command and updates arrive on their own:
/plugin marketplace add NirDiamant/Agentic_Engineering
/plugin install agentic-engineering@diamantai
/agent-memory
What it promises you:
- It writes
docs/, one instruction file at your repo root (plus a one-line pointer if your tools need the other name), a scorecard beside it, and the kit's two commands into.claude/commands/. Nothing else changes. - It deletes nothing of yours and overwrites nothing of yours. Your files are read as input. The most it does to one is append a dated entry to a log that is already append-only.
- It stops once, for your approval, before it writes the layer. The only thing on disk before that is your scorecard.
- About fifteen minutes.
Before it reads your code it answers five questions about your repo, cold. After the docs exist, it answers the same five again reading only what this run wrote, and prints the score, plus a count of the claims in your own docs that your code contradicts:
AGENT MEMORY CHECK: seo-autopilot
Before the docs layer: 3 of 5 questions right
After: 5 of 5
It was guessing about: which tools are installed vs only planned; the conventions the linter actually enforces; that packages/db is wired to nothing
Contradictions found in your existing docs: 5
Written: CLAUDE.md + 4 docs, 4 unconfirmed
Still advisory: nothing enforces these rules. That takes hooks, tests, CI.
A real run on a real repo, not a mock-up. It went 3 to 5, and the interesting half is the 3. That repo has a good README and a 600-line spec, so the agent had plenty to read. It read all of it and still named two tools that are installed nowhere in the tree, because they sit in the spec as a plan. It reported the plan as the stack, with the label SURE on it.
The contradictions line is the five claims in that repo's own README and spec that its code disproves: two tools named in the spec that no package installs, a scheduler the spec never chose doing the work, an upload the docs call unbuilt that shipped. Zero is a normal answer on a repo whose docs are current.
The run ends by offering a badge for your README, so the repo carries the score and the link back here:
If you run a team, the number you want is not one repository. Point this at the folder your repositories are cloned into and it reads each of them cold, the same five questions, and prints one card:
Clone https://github.com/NirDiamant/Agentic_Engineering into a temp folder, read its ORG.md, and follow it on this folder of repositories.
In Claude Code that is /org-card <folder>.
ORG MEMORY CHECK: orgtest
Repos read: 3
Scoring 2 or below: 3 of 3
Guessing most about: what naming or code convention a reviewer here would flag first
Instruction files: 0 of 3 repos have one
Contradictions on sight: 0
Rules enforced: 0 (nothing here checks a hook, a test or CI; that is the course)
A real run, on click, requests and fastapi cloned into one folder, 2026-09-22. All three are excellent libraries and all three scored 1 of 5, because a README tells an agent what the library is for and nothing about where a new file goes, why this dependency and not the obvious one, or what a change must never break. That is the point of the number: it is not a quality score, it is how much your agent is filling in from its training data.
That read writes nothing inside any repository. It is a measurement, and it runs on your machine, so no code leaves it. Full procedure: ORG.md.
One big repository that many teams share? A score for the whole monorepo measures everybody else's code too. Point it at the monorepo with your team's CODEOWNERS handle and it reads only your folders, each one as if it were a repository of its own:
/org-card . --owner @your-org/your-team
Without --owner it reads the packages your workspace manifest declares, and you can always name the folders yourself. On backstage, --owner @backstage/techdocs-maintainers found that team's 10 folders and scored 4 of the 10 at 2 or below, although every one of them inherits the repository's AGENTS.md. A root instruction file tells the agent how the monorepo works. It does not tell it where a new file goes in your plugin, or why your team picked the tool it did.
Each file answers one question your code can't. Here is TECH_STACK.md, filled in:
- **Database:** PostgreSQL
- **Why Postgres:** the team already knows it, and we don't need anything fancier.The second line is the one that matters. Your code already shows that you use Postgres. Why you chose it lives in someone's head, and an agent that never heard the reason will swap it for whatever it saw more often in training.
When the agent gets something wrong: "update CLAUDE.md so you don't repeat that, and tell me which old rule can come out."
When you stop for the day: "update docs/context/CONTINUE_PROMPT.md with where we are and what's next."
/start-project builds the same layer for a stack you've already chosen, plus SETUP.md, DEPLOYMENT.md, CI_CD.md and TESTING.md. /spark is for when it's still an idea: experimental, and it needs two skills that don't ship here, /harness-plan and /harness-unleash. /start-project and the other skills ship inside the plugin; /spark lives in skills-experimental/ and needs one extra step to install. The reasoning under all of it is in docs/why_this_works.md.
Your agent reads these files and tries to follow them. Nothing forces it. A rule that has to hold every single time belongs in a hook, a test, or a CI gate, and none of that ships here. Those layers are where the full method lives: Prompt to Production, my course, teaches them on your own project.
The final check is one model checking another. It catches empty placeholders and docs that contradict the code. It can't tell you the docs are true.
And it checks itself best in Claude Code. The after score is meant to be answered by a fresh reader, which means a sub-agent. Claude Code has them. Codex, Cursor and Copilot do not, so there the same session that wrote the files also grades them, and your card says so on its own line. Treat what you get in those tools as a first draft and read it before you trust it. A first draft of this carries real errors, and a confident wrong file is worse than no file.
Prompt to Production is my course on building software with AI the way professionals do: the methods and paradigms behind reliable, efficient, modular production systems, taught systematically. Every module pairs a video lecture with a hands-on lab, from your first structured prompt to a working production system. The hooks, tests and CI gates this kit leaves out are in there.
Your coding agent starts every session knowing nothing about your project, so it guesses. Paste one line into the agent you already have open, and about fifteen minutes later your repository has a docs layer written from the code itself, plus a card scoring what your agent knew before and after.
We ran it on six repositories you already depend on. Each one was asked five questions about itself, cold, then again after the layer was written. The last column counts statements in that project's own documentation that its own code disproves:
| repo | before | after | own docs its code disproves |
|---|---|---|---|
| fastapi | 2 of 5 | 5 of 5 | 2 |
| flask | 3 of 5 | 5 of 5 | 6 |
| django | 3 of 5 | 4 of 5 | 1 |
| express | 2 of 5 | 4 of 5 | 4 |
| requests | 2 of 5 | 4 of 5 | 0 |
| langchain | 5 of 5 | 4 of 5 | 5 |
Flask's six include four documentation examples that raise TypeError when you run them. Langchain scored lower afterwards, because it already ships a 380-line agent instruction file and the cold read was grading theirs; that row is in the table anyway.
Clone any of those repos, paste the same line, and check the number yourself. No signup.
| ๐ Weekly Updates |
๐ก Expert Insights |
๐ฏ Top 0.1% Content |
Join 40,000+ readers getting clear AI tutorials every week.
Contributing. Sharpen /apply or tighten a template. Issues and PRs are welcome. Apache 2.0, see LICENSE, and copy the output into any project, commercial ones included. These are starting points and not guarantees. They don't replace tests, review, or judgment.
More open-source guides: RAG ยท Agents ยท Prompting ยท Agent memory
If this saved you a session of re-explaining your project, a star helps other people find it.

