tonydzi/deep-research
226 deep-research reports from Palo Alto AI Research Lab, distilled. Why we ran each one, and what we got. Multi-model methodology: every question fanned out to several frontier models, then reconciled. CC BY 4.0.
Technical lead, Palo Alto AI Research Lab. Security engineer (MSc, cryptography) — evidence-first reliability work across the AI-agent ecosystem.
226 deep-research reports from Palo Alto AI Research Lab, distilled. Why we ran each one, and what we got. Multi-model methodology: every question fanned out to several frontier models, then reconciled. CC BY 4.0.
One person, six packagings - CV variants generated from a single resume.json master
START HERE — Anton Dziatkovskii, one page of proof
Founder, Palo Alto AI Research Lab — evidence-first reliability work across the AI-agent open-source ecosystem
C(H+A)RM build diary — a CRM for Human & Agent collaboration, built in public with Claude Code. Framework: github.com/tonydzi/charm-os. English diary, longreads & reusable artifacts. RU: Telegram @ClawRus.
Graph RAG on SQLite for AI agents: vector retrieval + hand-curated wikilink graph + cross-encoder rerank, with a zero-token per-turn memory ledger. Working pilot.
Claude Code as a second brain: 100 battle-tested skills, a working CRM engine, vault templates and the handover map. Method, not data. MIT.
Lossless storage and search for AI agent sessions, across every agentic client.
Chrome DevTools for coding agents
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Run AI agents across machines without state drift: consensus protocol (propose/counter/accept/commit), dual-rail message bus, ACK discipline, self-healing sync. Formerly claude-consensus. Free, MIT.
Make Claude Code behave consistently across repos, sessions and humans: rules-as-files governance with precedence, declined-decisions journal, objection sparring. Sanitized from a live production system. Free, MIT.
The Journey — a build-in-public book of a 60-day founder+AI collaboration. Two voices (Tony & Mycroft), human + machine readable. Palo Alto AI Research Lab.
Link WhatsApp to Claude (or any MCP client) in ~20 minutes: a live self-refreshing QR page that makes pairing actually work, production patches, and 12 field gotchas. On top of @sjawhar/whatsapp-mcp.
Voice -> text -> your personal knowledge base. A primitive, not a platform: 4 small Python scripts turn voice notes into linked, tagged, searchable markdown.
Daily lessons: Claude Code and Codex for people who have never opened a terminal. For business owners, not programmers. Written from a running fleet of AI agents, not from theory.
Your scheduled job says exit 0 — prove it did the work. Three stdlib-only checks: output freshness, silent no-op detection, rollout proof by reading the fact back. Sanitized from a live agent fleet. Free, MIT.
Your LLM reviewer said APPROVE. Did it? A structured verdict contract: prompt rule + parser + exit-code gate in one stdlib file, with the 42 counterexamples that forced every line. Free, MIT.
Connect Claude (or any MCP client) to your own Telegram in ~15 minutes: setup prompt for Claude Code/Codex, production patches (shared daemon, extra tools, multi-account), watchdog, and every gotcha that cost us hours. On top of chigwell/telegram-mcp.
Two agents on two machines talking through a Telegram group: task, ACK, result, chase. No ports, no VPN, no server. Telegram bots cannot see each other, so this uses user sessions - the reason nobody else has built this. stdlib, MIT.
Nobody reviews themselves - and one reviewer model is one blind spot. Fan every change out to several model FAMILIES at once: verdict contract, quorum by family, honest skips, and an exit code that refuses to go green on one vendor. Stdlib-only, MIT.
Review a PR by measuring which of its own guards its own test suite actually catches. Neutralise each guard, run their tests, report the removals nothing notices. Stdlib-only, no install. Claude Code skill + standalone tool.
One persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.
Outbound ledger + daily harvest digest for your PRs and issues in other people's repos. One file, stdlib, gh CLI, zero LLM.
Open up your internal work without leaking it: substitute personal data with plausible fakes of the same shape (not <REDACTED>), then gate the WHOLE tree before the push. Stdlib-only, with the eleven traps it cost us.
One question, nine independent paid LLM accounts in their strongest reasoning mode, and a published breakdown of where they DISAGREE. Failures included by rule.
One shared MCP daemon per machine instead of a stdio copy in every agent session: recipe, autostart templates for Windows/macOS/Linux, a watchdog that will not blind your live sessions, and the measurements to prove it
Your agent setup charges rent on every session, before it does any work. Three stdlib-only instruments: what your wiring costs per session, who burned yesterday, and which paid subscriptions are going undrawn. 0 LLM, 0 network, MIT.
MCP server for the Lambda Cloud API: GPU prices, live capacity, cheapest-available lookup, and env-gated launch/terminate with a price cap. Read-only by default.
Pre-registered A/B/C benchmark: Superpowers vs OpenSpec vs baseline, on real brownfield ops code. Methodology frozen before runs; results published either way.