jimbobbennett/mastering-ai-agents-tracing-evals

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

Understand, test, and fix agent failures with tracing and evals

A hands-on workshop notebook on finding the failures in an AI agent that don't throw exceptions, don't look wrong in the output, and never show up in your logs.

Presented by Jim Bennett.

Open In Colab

What you'll build

A shopping assistant for a toy store called Wonder Toys — a LangChain agent with five tools that searches a 200-product catalogue, looks up products, and takes orders.

It works. It answers well. It also has a bug that costs you an extra tool call and an extra model round-trip on a large slice of requests, and you cannot see it from the answers.

What you'll learn

Step 1 — Build the agent. Tools, a system prompt, semantic search over a local ChromaDB, and conversation memory. Run it and watch it give good answers.

Step 2 — Add tracing. Wire up Arize AX with OpenTelemetry and OpenInference in about three lines, without touching the agent. Then answer the questions about your own agent that you couldn't answer a moment earlier: how many model calls, how many tool calls, with what arguments, at what cost.

Step 3 — Find the bug. Read the traces and spot the failure. Learn what an agent's trajectory is, why convergence is the thing that's wrong here, and where human review fits in.

Step 4 — Write an eval. Build an LLM-as-a-judge evaluator that detects the failure as a shape rather than a specific case, so it generalises to tools you haven't written yet. Run it over your traces, read the scores back, and ask the awkward question of who judges the judge.

Step 5 — Fix it and prove the fix. Repair the tool, re-run the same prompts, and let the eval tell you whether it worked. Test-driven development, with an eval as the test.

Running it

Easiest is Colab — click the badge above, no local setup.

To run locally, clone the repo and open mastering-ai-agents.ipynb in Jupyter or VS Code. Every dependency installs from inside the notebook.

You'll need:

  • An OpenAI API key — platform.openai.com/api-keys. The whole notebook costs well under a dollar to run.
  • An Arize AX account — free to start. You need your Space ID and an API key, both from Settings → API keys.

Two things worth knowing before you start. The first search cell downloads an embedding model of about 80 MB and caches it, so run the setup cells early if you're on shared conference wifi. And traces take a few tens of seconds to become queryable in Arize, so an empty project straight after a run isn't a problem yet.

What's in the repo

mastering-ai-agents.ipynb The workshop
products.json The 200-product toy catalogue
img/ Screenshots used in the notebook

Further reading

Contributors

jimbobbennett

Issues