← Course overview

Module 13 of 14

Production Agent Systems

What changes when an agent moves from a demo to production?

Learning objectives. Move a system from prototype through pilot to production; instrument observability; manage cost, versioning and maintenance; design escalation and incident response; and drive adoption.

Core lesson

Three stages, three different jobs

PrototypePilotProduction
PurposeDoes this work at all?Does it work for real users on real work?Does it work reliably, affordably, forever?
UsersYouA named small groupEveryone in scope
DataSamplesReal, with supervisionReal, at volume
SuccessAn impressive runMeasured improvement with a manageable exception rateStable metrics, controlled cost, low maintenance
Typical durationDaysWeeksYears

The gap between prototype and production is not ten per cent more work; it is most of the work. Everything in this module is that gap.

ARCHITECT — observability

You cannot operate what you cannot see. Instrument four things:

Traces. For every run: the objective, every model call with its prompt version, every tool call with arguments and results, every decision point, the final outcome, and the total cost. A trace must let a person reconstruct exactly why the agent did what it did, months later.

Metrics. Completion rate, error rate by type, tool success rate, latency percentiles, cost per run and per successful task, human intervention rate, and the business KPI. Track them per version.

Logs. Structured and searchable, with correlation IDs linking a user request to every downstream call.

Alerts. On error-rate spikes, cost anomalies, latency regressions, unusual tool-use patterns, spikes in escalation, and any authorisation failure — the last of which is often the first sign of an attack.

Cost management

Costs in agent systems grow quietly. Control them deliberately: route to the smallest sufficient model; cache aggressively — identical prompt prefixes, repeated retrievals, deterministic sub-results; prune context; cap steps per run; run batch work on cheaper tiers; and set hard spend ceilings per agent, per run and per day, with alerting well below the ceiling. Report cost per successful task to the business; it is the only figure that connects spend to value.

Versioning and change management

Version prompts, tool schemas, models, retrieval indexes, policies and evaluation sets. Every change runs the regression suite. Roll out significant changes progressively — a small percentage of traffic first — and keep the ability to roll back within minutes. Pin model versions where the provider allows it: a silently updated model is an unversioned change to your system.

Failure handling and escalation in production

Design the degraded modes explicitly: fall back to another model on provider failure; fall back to a simpler deterministic path; queue rather than fail; and hand to a human with full context. Users should experience “this one is going to a specialist”, never a stack trace. Every escalation should carry what was attempted, what was learned and what remains — an escalation is a hand-over, not an abandonment.

Incident response for agent systems needs AI-specific playbooks: an agent taking incorrect actions at volume, a prompt-injection incident, a cost runaway, a provider outage, a data-exposure event. Each needs a detection signal, a kill switch, a containment step, a communication plan and a post-incident review that feeds the evaluation set.

Maintenance — the line everyone forgets

Agent systems decay. Models are deprecated; APIs change; documents go stale; business processes shift; edge cases accumulate. Budget twenty to thirty per cent of build effort annually for maintenance and name an owner. An agent without a named owner is an incident waiting to happen.

Adoption

More agent projects die of non-adoption than of technical failure. What works: solve a problem the users actually complain about; involve them in defining success; be explicit about what the agent does and does not do; make the review step fast and pleasant; show the time saved; make it trivially easy to report a bad output; and act visibly on those reports. What fails: launching to everyone at once, hiding the AI, over-promising, and treating scepticism as resistance rather than as unpaid quality assurance.

Business example

A firm ran a successful three-month pilot with four users and rolled the agent out to forty. Within two weeks, cost was triple the forecast (nobody had pruned context for high-volume users), the escalation queue was unstaffed, and three users had quietly reverted to the manual process. The recovery was not technical: a spend ceiling with alerting, a named queue owner, a fifteen-minute onboarding session per team, and a weekly published scorecard. Adoption reached ninety per cent in six weeks. Production is an operational discipline, not a deployment event.

Common mistakes

  • Treating deployment as the finish line.
  • No tracing, so failures are irreproducible.
  • No spend ceiling.
  • Never pinning model versions.
  • Escalation with no owner.
  • No maintenance budget, so quality erodes invisibly across a year.
  • Rolling out to everyone simultaneously.

Expert insight

The teams that succeed in production are rarely the ones with the most sophisticated architecture. They are the ones with tracing, a regression suite, a spend ceiling, a named owner and a weekly scorecard. These are unglamorous, and they are the difference between a system that compounds value and one that quietly becomes a liability.

Knowledge check

  1. Distinguish prototype, pilot and production by purpose and success criteria.
  2. What must a trace contain to be useful in an investigation?
  3. Name five cost-control levers.
  4. Why must model versions be pinned?
  5. What belongs in an AI-specific incident playbook?
  6. What annual maintenance budget should be assumed, and why?
  7. Name three practices that drive adoption and three that destroy it.

Challenge

Write a one-page operations runbook for an agent: owner, kill switch, alert thresholds, escalation path with a named queue owner, spend ceiling, rollback procedure, and the monthly review agenda. If you cannot complete it, the agent is not ready for production.