← Course overview

Module 12 of 14

Evaluating AI Agents

How do you know if an agent is actually working?

Learning objectives. Build an evaluation set; choose metrics that reflect business reality; measure reliability across repeated runs; use LLM-as-judge appropriately and know its limits; and run regression testing so improvements do not silently break things.

If you cannot measure an agent, you cannot reliably improve it — and you should not deploy it.

Core lesson

Why single-run testing is worthless

Agents are non-deterministic. A run that succeeds tells you the task is possible, not that the system is reliable. The distinction has a formal expression worth adopting.

pass@k asks: across k attempts, did at least one succeed? Useful for research capability claims.

pass^k asks: across k independent attempts, did every one succeed? This is the number that matters in production, because your client does not get to retry until it works. An agent succeeding eight times in ten looks strong under pass@k and fails badly under pass^k. Adopt pass^k, or a straightforward consistency rate over repeated runs, as your reliability headline. It is usually a sobering number, and that is the point.

The evaluation set

Twenty to fifty representative tasks with known-good outcomes will find most of your problems; you do not need thousands to start. Compose it deliberately:

  • Typical cases (roughly half) — the everyday work
  • Edge cases (roughly a quarter) — unusual but legitimate
  • Failure cases (roughly a quarter) — inputs where the correct behaviour is to refuse, ask, or escalate

That last category is the one teams omit, and it is the one that measures judgement. An agent that never says “I need more information” is not confident; it is unmeasured.

Freeze the set. Version it. Add every production failure to it permanently — this is how an evaluation suite becomes an asset rather than a chore.

PRACTITIONER — the metric stack

DimensionMetricNotes
Task completion% of tasks completed to specificationThe headline capability number
Accuracy% of outputs factually correct against ground truthRequires ground truth; sample and audit where it does not exist
Reliabilitypass^k / consistency across repeated runsThe production number
Tool success% of tool calls that are well-formed and succeedIsolates integration problems from reasoning problems
Retrieval quality% of queries where the correct source is retrievedEvaluate separately — it caps everything downstream
Grounding% of factual claims traceable to a cited sourceThe practical hallucination measure
LatencyMedian and 95th percentile end-to-endThe p95 is what users remember
CostCost per successful taskNever cost per token
Safety% of unsafe or out-of-scope requests correctly refusedInclude adversarial cases
Human interventionOverride rate, approval rate, escalation rateVery high means not ready; near-zero often means rubber-stamping
Business outcomeThe KPI the agent exists to moveThe only one that ultimately justifies the system

Evaluate layer by layer. End-to-end scores tell you that something is wrong. Component evaluation — retrieval, tool calls, individual agents, the synthesis step — tells you what.

LLM-as-judge: useful, and limited

Using a model to grade outputs against a rubric scales evaluation enormously and is now standard practice. Anthropic’s published account of evaluating its research system describes exactly this: a single judge model scoring against a rubric covering factual accuracy, citation accuracy, completeness and source quality, alongside human testing that catches what automation misses.

Use it well: a specific rubric with defined levels, not “rate this 1–10”; a judge model that is not the generator; a held-out set of human-graded examples to calibrate the judge itself; and a periodic human audit of a sample. Know its limits: judges are biased toward verbose and confident output, they share the generator’s blind spots when they share its training and prompt, and they cannot assess business appropriateness. The Anthropic account is instructive on this point — human testers, not automation, discovered that the agents systematically preferred SEO-optimised content farms to authoritative sources.

Regression testing

Every change to a prompt, a model, a tool schema, a retrieval index or a policy is a change to system behaviour. Run the full suite before shipping. Record scores per version. Do not ship a change that improves the target metric and degrades another without an explicit, documented decision. Without this discipline, agent systems drift downwards in ways nobody can explain months later.

The agent scorecard

A one-page artefact, published per agent, reviewed monthly:

FieldExample
Agent, version, ownerReporting Agent v2.3 — Head of Delivery
MissionOne sentence
Evaluation set40 tasks, v6, last updated 12 Aug
Task completion91%
Reliability (pass^5)78%
Grounding96% of claims cited
Cost per successful task£0.42
p95 latency3m 10s
Human override rate12%
Business KPIPrep time 90 to 25 min; on-time delivery 68% to 97%
Open issuesUnder-detects seasonality in Q4 comparisons
DecisionContinue; retrain rubric on seasonality cases

Business example

A customer-service drafting agent reported ninety-four per cent accuracy and was approved. In production, satisfaction fell. Investigation found the evaluation set contained only typical enquiries; the real inbox was a third complaints and edge cases, where the agent was fluent and wrong. The rebuilt evaluation set was drawn from actual traffic, including the messy quarter. Measured accuracy dropped to seventy-one per cent — a truer number — and the redesign that followed routed complaints straight to humans and applied the agent only where it was genuinely strong. Composite accuracy was the wrong metric; accuracy per category was the right one.

Common mistakes

  • Testing on cases you designed for, and calling it evaluation.
  • Single-run testing.
  • Cost per token instead of cost per successful task.
  • No refusal cases, so judgement is never measured.
  • LLM-as-judge with no human calibration.
  • Never adding production failures to the suite.
  • Reporting a composite score that hides a catastrophic sub-category.

Expert insight

Evaluation is the discipline that converts opinion into engineering. Teams without it argue about whether the agent “feels better”; teams with it ship measured improvements and catch regressions before clients do. Build the evaluation set before you build the agent — it is also, unexpectedly, the clearest specification of what you are building.

Knowledge check

  1. Distinguish pass@k and pass^k, and say which belongs on a production scorecard.
  2. What three categories should an evaluation set contain, and in roughly what proportions?
  3. Why must retrieval be evaluated separately?
  4. Give three limitations of LLM-as-judge.
  5. Why is a near-zero human override rate potentially a warning sign?
  6. Why is cost per successful task the correct cost metric?
  7. What should happen to every production failure?

Challenge

Build a thirty-task evaluation set for an agent you use, including at least eight cases where the correct behaviour is to refuse or escalate. Run each task five times. Report task completion, pass^5, and cost per successful task. Most teams find their reliability is materially below their assumption.

Further exploration

τ-bench (tool–agent–user interaction with pass^k scoring), SWE-bench Verified, and GAIA as reference points for how the field measures agents — useful for calibration, never a substitute for evaluation on your tasks.