The Holistic Agent Canvas
A structured, 20-element framework for designing an AI agent's objective, permissions, tools, memory, and governance before a single line of code is written — vendor-neutral and free to use.
The HolisticAgent.ai framework — Five Planes, Twenty Elements (HAC-20)
Every agent that works in production, whatever the vendor, framework or model, resolves the same twenty questions. Most failed agents failed because somebody never asked one of them.
The Holistic Agent Canvas organises those twenty elements into five planes. It is a design tool, a review checklist and an audit artefact. You fill one in before you build, you revise it as you learn, and you attach it to the system as documentation.
Figure 2 — The Holistic Agent Canvas: five planes, twenty elements
Plane 1 — PURPOSE: what this agent is for
| Element | The question it answers | Failure if unanswered |
|---|---|---|
| 1. Identity | Who or what is this agent? What is its name, role, voice and standing with users? | Users cannot calibrate trust; the agent’s tone is inconsistent with the brand |
| 2. Mission | Which single outcome is it accountable for? Stated as an outcome, not an activity. | Scope sprawl; nobody can say whether it succeeded |
| 3. Boundaries | What is explicitly out of scope? What must it refuse or escalate? | The agent attempts work it is unfit for, confidently |
A well-formed mission reads: “Reduce the time to produce a first-draft technical SEO audit from four hours to under thirty minutes, at equal or better accuracy, with a human reviewer approving before client delivery.” An ill-formed mission reads: “Use AI for SEO audits.”
Plane 2 — COGNITION: how it understands and decides
| Element | The question it answers |
|---|---|
| 4. Context | What must it know at this moment to act correctly? Instructions, state, current inputs, recent history. |
| 5. Knowledge | Which corpora, documents, records and systems of record can it consult? |
| 6. Memory | What persists between runs, for how long, and who can correct it? |
| 7. Reasoning | How does it determine what to do — direct response, chain of thought, extended reasoning, critique? |
| 8. Planning | How does it decompose an objective into steps, and how does it re-plan when a step fails? |
Plane 3 — ACTION: how it changes the world
| Element | The question it answers |
|---|---|
| 9. Tools | Which capabilities can it invoke, with what inputs, and what does each one cost? |
| 10. Environment | Which systems, data stores and interfaces does it operate inside? Sandbox, staging or production? |
| 11. Action | What can it actually change? Which of those changes are reversible, and which are not? |
| 12. Observation | How does it learn what actually happened — return values, errors, state checks, verification steps? |
The reversibility question in element 11 is the single highest-leverage design question in the whole canvas. Sort every action into three buckets — reversible, reversible with effort, irreversible — and let that sorting, not enthusiasm, determine how much autonomy you grant.
Plane 4 — GOVERNANCE: the boundaries autonomy requires
| Element | The question it answers |
|---|---|
| 13. Permissions & Identity | Under whose identity does the agent act? What is the least privilege that still allows the mission? |
| 14. Policy | Which business rules, legal constraints and brand standards must it obey? How are they enforced — in the prompt, in code, or in the platform? |
| 15. Human Oversight | Where exactly does a human approve, review, sample or intervene? Named role, not “someone”. |
| 16. Security & Audit | How is it protected from prompt injection, tool poisoning, data leakage and privilege abuse? What is logged, and for how long? |
Enforce policy in the strongest available layer. A rule written only into a system prompt is a request. The same rule enforced by a permission scope, a schema constraint or a code check is a control. Auditors and attackers both know the difference.
Plane 5 — EVOLUTION: how it gets better and stays affordable
| Element | The question it answers |
|---|---|
| 17. Evaluation | How do you know it worked? Which test set, which metrics, which threshold to ship? |
| 18. Feedback | How do corrections, complaints and human edits flow back into improvement? |
| 19. Economics | What does a run cost in tokens, tool calls, latency and human review minutes? What is it worth? |
| 20. Reliability & Evolution | How consistently does it perform across repeated attempts, and how is it versioned and maintained over time? |
Using the canvas
START HERE. Print the canvas. For any process you are tempted to automate, fill in Plane 1 and Plane 4 first. If you cannot state the mission in one sentence or name the human who is accountable, do not build yet.
PRACTITIONER. Fill all five planes before writing a single prompt. Most of your build time will be spent on Planes 2 and 3; most of your production incidents will come from Planes 4 and 5.
ARCHITECT. Treat the canvas as a specification artefact under version control. Each element maps to something testable: mission to success criteria; tools to contract tests; permissions to an access review; economics to a cost-per-successful-task metric; reliability to a pass^k measurement (Module 12).
The Holistic Agent Loop
The canvas describes the anatomy. The loop describes the physiology — what actually happens each time the agent runs.
Figure 3 — The Holistic Agent Loop
Read it as a sentence: an objective is set; context is assembled; the agent reasons; it forms or revises a plan; it takes one action through a tool; it observes the real result; it evaluates progress against the objective; and it either loops, escalates to a human, or stops.
Three properties of this loop separate agents from automations:
- The number of iterations is not known in advance. This is the defining technical characteristic of an agent, and the source of most of its cost and risk.
- The loop is grounded by observation. An agent that never checks the result of its actions is not an agent; it is a text generator with side effects.
- The loop must have terminators. Step budgets, cost ceilings, time limits, confidence thresholds and escalation triggers are not optional. An agent loop without a terminator is a defect waiting for an invoice.
The Agency Test
Before how do I build an agent? comes should this be an agent at all? Run the six questions below. Each “no” pushes you down the conceptual ladder toward something simpler, cheaper and more reliable.
| # | Question | If NO, use instead |
|---|---|---|
| 1 | Is the sequence of steps genuinely unpredictable in advance? | A deterministic workflow or script |
| 2 | Does the task require judgement over unstructured information? | A database query, report or rules engine |
| 3 | Is there a real tolerance for variance in how the task is done? | A fixed automation |
| 4 | Can success be checked — by code, by a rubric, or by a human? | Do not automate yet; define success first |
| 5 | Is the value of a successful run comfortably above its cost? | A cheaper method, or a human |
| 6 | Are the consequences of a wrong action bounded and recoverable? | Human-in-the-loop, or do not automate |
The Agency Test in one line: use an agent when the path is unknown, the judgement is real, success is checkable, the economics work and the blast radius is bounded.
Most work in most businesses fails at least one of these questions. That is not a disappointing finding; it is the finding that saves budgets. The mature practitioner deploys deterministic automation for the eighty per cent of work that is genuinely predictable, and reserves agents for the twenty per cent that genuinely is not.