Module 10 of 14
Agent Architecture and Orchestration
How do the pieces fit together at a system level?
Learning objectives. Describe the reference architecture of a production agent system; choose between loop, state machine and event-driven designs; implement model routing, state persistence and failure recovery; and design long-running agents that survive restarts.
Core lesson
The reference architecture
Figure 6 — Reference architecture of a production agent system
Every serious agent system contains the same eight layers, whether you build them or a platform provides them:
| Layer | Responsibility |
|---|---|
| Interface | How work arrives: chat, form, email, webhook, schedule, event |
| Orchestration | Routing, decomposition, delegation, loop control, terminators |
| Reasoning | Model calls, model selection, prompt and context assembly |
| Context & memory | Retrieval, state, compaction, note-taking, persistence |
| Tools | Registry, schemas, execution, rate limits, error handling |
| Governance | Identity, permissions, policy enforcement, approvals, audit |
| Observability | Tracing, logging, metrics, cost accounting, alerting |
| Storage | State, artefacts, memories, evaluation data, versions |
If you cannot say where each of these lives in your system, you have a prototype rather than an architecture.
Loop, state machine, or event-driven?
The agent loop. The model decides the next action each iteration. Maximum flexibility, minimum predictability. Correct when the path is genuinely unknown. Requires strict terminators.
The state machine (structured workflow). Defined states with defined transitions; agentic reasoning happens inside specific states. Far more testable, auditable and cheaper. This is the right default for most business processes, and the honest description of most successful production “agents”.
Event-driven. Agents wake on events — a message arrives, a file lands, a threshold trips, a schedule fires. Natural for operational work, decouples components, and demands idempotency and deduplication because events arrive twice.
Most real systems are hybrids: an event triggers a state machine, and two of its states contain bounded agent loops. Use the least dynamic structure that solves the problem.
ARCHITECT — the hard parts
Model routing. Classify the step, then choose the model: a small fast model for classification, routing and extraction; a mid-tier model for the bulk of drafting and analysis; a frontier model for genuinely hard reasoning; a fallback model for provider outages. Route on task type, not on prestige. This is typically the single largest cost lever in a mature system, frequently reducing spend by more than half with no measurable quality loss — but only if you have an evaluation set to prove the “no measurable quality loss” part.
State persistence. Long-running agents must survive restarts, deployments and context exhaustion. Persist: the objective and acceptance criteria; the plan and its progress; established facts and decisions; completed actions with their results (for idempotency); open questions; and the cost consumed so far. A durable, human-readable progress file is both an engineering mechanism and a governance artefact: the agent’s state becomes something a person can read, correct and hand over.
Checkpoints and recovery. Take a checkpoint at every meaningful state transition. On failure, resume from the last good checkpoint rather than restarting; restarting a long agent from zero is both expensive and, where writes occurred, dangerous. Version-controlled artefacts serve the same purpose for work products: they let the system revert a bad change and recover a known-good state.
Asynchronous and long-running work. Work measured in minutes or hours needs queues, a task record with a status the user can poll, timeouts with defined behaviour, and progress reporting. Never hold a user-facing request open for a long agent run.
Concurrency. Parallel agents writing to shared state need locking or ownership rules. The commonest production bug in multi-agent systems is two agents updating the same record with divergent conclusions.
Failure recovery taxonomy. Transient (retry with backoff); permanent (fallback or escalate); partial (roll back or compensate — and know which of your actions can be compensated); ambiguous, where you do not know whether the write succeeded (check state before retrying, always).
Versioning. Prompts, tool schemas, models, retrieval indexes and policies are all versioned artefacts. A change to any of them is a change to system behaviour and must trigger regression evaluation. Teams that version code but not prompts will experience silent, unexplainable behaviour changes.
Business example
A document-processing system handled ten thousand documents a month. The first architecture was a single agent loop per document: unpredictable cost, occasional runaway loops, no visibility. The second was a state machine — receive › classify › extract › validate › route › archive — with an agent operating only inside classify and extract, a small model doing classification, a mid-tier model doing extraction against a strict schema, deterministic validation, and human review only for documents failing validation. Cost per document fell by roughly seventy per cent, throughput became predictable, and every document had a traceable state history. The reduction in autonomy was the improvement.
Common mistakes
- Building an open-ended loop where a state machine was obvious.
- One model for everything.
- No state persistence, so any interruption destroys hours of work.
- Retrying a write without checking whether it already succeeded.
- Unversioned prompts.
- No trace. If you cannot reconstruct why the agent did what it did, you cannot operate it (Module 13).
Expert insight
Architectural maturity in this field looks like reducing dynamism over time. Teams begin with maximal agency because it demos well, then discover that most of their process is knowable, and progressively convert agentic steps into deterministic ones — keeping agency only where variance genuinely lives. The end state is not less capable; it is cheaper, faster, more reliable and explainable. Agency is expensive; spend it where it earns its keep.
Knowledge check
- Name the eight architectural layers and one responsibility of each.
- When is a state machine preferable to an open agent loop?
- What must be persisted for a long-running agent to survive a restart?
- Why must event-driven agents be idempotent?
- Describe the four failure categories and the correct response to each.
- Why is model routing usually the largest cost lever?
- Why must prompts be versioned as strictly as code?
Challenge
Draw the reference architecture for one agent you are building, naming the concrete component at every layer. Identify the two layers you have not yet built. They are almost always governance and observability — and they are the two that determine whether the system reaches production.