← Course overview

Module 7 of 14

Planning, Reasoning and Decision Making

How does an agent decide what to do next across many steps?

Learning objectives. Distinguish thinking, planning, executing and verifying; select reasoning strategies proportionate to task difficulty; design retries, fallbacks and escalation that do not amplify failure; and set decision thresholds that route work to humans appropriately.

Core lesson

Four distinct activities

Thinking is producing intermediate reasoning to reach a better conclusion in one step. Planning is decomposing an objective into ordered steps with dependencies. Executing is taking an action through a tool. Verifying is confirming the action achieved what was intended. Teams routinely build agents that think and execute, and simply omit planning and verification. Those agents are fast, confident and unreliable.

PRACTITIONER — reasoning strategies and when to use them

StrategyWhat it doesUse whenAvoid when
Direct responseAnswer in one stepSimple, well-defined tasksMulti-step or high-stakes work
Chain of thought / extended reasoningReason before answeringMulti-step deduction, analysis, ambiguityClassification and extraction — it adds cost without accuracy
DecompositionBreak into sub-tasks explicitlyComplex objectives with distinct partsSimple tasks; it adds coordination overhead
Self-critique / reflectionReview own output against criteriaQuality-sensitive output; draftingThe critic shares the generator’s blind spots — prefer an external check where one exists
Verification by toolCheck the claim against realityAnything checkable: numbers, existence, stateGenuinely subjective judgements
Voting / samplingRun several times and compareHigh-stakes, variance-sensitive decisionsCost-sensitive routine work

The verification hierarchy, strongest first: check against ground truth in a system (did the record actually change?); check against a deterministic rule (does it balance?); check with an independent model or agent; ask the model to check itself; assume it is fine. Most production incidents in agent systems come from teams operating three or four levels below the strongest verification available to them.

Sequential versus parallel. Steps with dependencies must be sequential. Independent steps — researching five competitors, checking six pages, drafting three variants — should be parallel: it collapses latency, and each branch keeps a clean context. Parallelism costs tokens and complicates error handling; use it where the work genuinely is independent.

Re-planning. A plan made before the first tool call is a hypothesis. Real environments invalidate it: the record does not exist, the format differs, the API is down. Agents need an explicit re-planning trigger — a failed step, a contradicted assumption, a step budget threshold — otherwise they either abandon the objective or keep executing an invalid plan.

ARCHITECT — failure handling that does not amplify

Retries must be bounded, backed off, and different. Retrying an identical failing call three times is superstition; retrying with a corrected argument, a different tool, or a narrower scope is recovery. Never retry a non-idempotent write without a check that the first attempt did not succeed.

Fallbacks form a chain: preferred tool › alternate tool › cached or degraded answer › human escalation. Each level must be explicitly designed, including what the user is told.

Uncertainty and decision thresholds. Models are poorly calibrated: stated confidence does not reliably track accuracy. Build thresholds on observable signals instead — retrieval scores, agreement between independent runs, schema validation results, whether a verification tool confirmed the state, whether the value at risk exceeds a limit. Then route: proceed autonomously; proceed and notify; require approval; or refuse and escalate.

Loop control. Every loop needs a step budget, a spend ceiling, a wall-clock limit and a no-progress detector — if the last three iterations produced no state change, stop. On termination, preserve state and escalate with a summary of what was attempted, what was learned, and what remains. A well-designed escalation is a deliverable, not an error message.

Business example

A finance reconciliation agent matched transactions against invoices. Version one reasoned about each transaction and produced a confident narrative — which was wrong roughly one time in twelve, invisibly. Version two changed almost nothing about the reasoning and added one thing: after proposing a match, the agent called a tool that verified the amounts, dates and references actually agreed, and every unverified match was routed to a human queue. Error rate on accepted matches fell to near zero; roughly nine per cent of items went to the human queue. Verification converted an unreliable agent into a reliable system with a manageable exception rate.

Common mistakes

  • Applying maximum reasoning to every task, tripling cost for no accuracy gain.
  • Self-critique as the only quality control, when an external check exists.
  • Unbounded retries producing runaway cost and duplicate side effects.
  • No re-planning trigger, so the agent executes an invalidated plan to completion.
  • Treating the model’s stated confidence as a calibrated probability.
  • Parallelising dependent steps and producing incoherent results.

Expert insight

Reliability comes from verification, not from cleverness. Given a fixed effort budget on an underperforming agent, invest it in checking outcomes before you invest it in improving reasoning. The strongest agent architectures are ordinary reasoning wrapped in rigorous verification.

Knowledge check

  1. Distinguish thinking, planning, executing and verifying.
  2. When does extended reasoning not help?
  3. State the verification hierarchy from strongest to weakest.
  4. What makes a retry legitimate rather than superstitious?
  5. Why should decision thresholds not be based on the model’s stated confidence?
  6. Name four loop terminators and describe what should happen at termination.
  7. Why is a good escalation a deliverable?

Challenge

Take an agent task that fails occasionally. Add exactly one verification step at the strongest available level in the hierarchy, and route unverified results to a human queue. Measure the error rate on accepted results and the size of the exception queue before and after. Report the trade-off you have chosen.

Further exploration

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, arXiv:2210.03629 — the foundational reason–act–observe loop.