Module 4 of 14
Prompting Is Not Enough
Why does reliable agent behaviour depend on context, not clever wording?
Learning objectives. Explain the progression from prompt engineering to context engineering to agent engineering; assemble a context deliberately rather than accumulate one; write tool descriptions and system instructions that hold up under load; and know why longer prompts do not produce better agents.
Core lesson
The three eras
Prompt engineering asks: what do I say to the model to get a good answer? It optimises a single string for a single turn. It remains a genuine skill and it is not sufficient.
Context engineering asks: what is the smallest, highest-signal set of tokens that should be in the model’s view at this step? The unit of design is no longer a sentence but a state — the full assembly of system instructions, task instructions, retrieved knowledge, prior results, tool definitions and tool outputs. This is a curation problem, and curation means exclusion.
Agent engineering asks: what is the system that repeatedly assembles the right context, chooses the right action, verifies the result, controls its cost and knows when to stop? The unit of design is the loop and everything around it.
Figure 4 — From prompt engineering to agent engineering
PRACTITIONER — the anatomy of a well-built context
| Layer | Contains | Design rule |
|---|---|---|
| System instructions | Identity, mission, boundaries, tone, non-negotiable rules | Short, unambiguous, in priority order. If it is enforceable in code, enforce it in code instead. |
| Task instructions | The current objective and its acceptance criteria | Concrete and checkable |
| Structured knowledge | Retrieved documents, records, policies | Just-in-time, cited, minimal |
| Working state | What has been established, decided or completed so far | Externalised to a file or store; summarised into context, not pasted whole |
| Tool definitions | Names, descriptions, schemas | Few, distinct, unambiguous |
| Tool results | Observations from the environment | Filtered before entering context; clear old results aggressively |
| Output contract | The exact shape of the expected result | A schema beats a description |
Structured outputs deserve emphasis. Asking a model to “reply in JSON with these fields” is a request. Constraining generation to a schema is a guarantee of shape. It eliminates a whole class of parsing failures and makes downstream code simple and testable. Wherever the next step is code, use a schema.
Tool descriptions are prompts. They are the most under-invested text in most agent systems. A tool description should state what it does, when to use it, when not to use it, what it costs if that is significant, and what it returns. Tools whose descriptions overlap will be chosen inconsistently.
Examples beat adjectives. Two well-chosen examples of correct output do more than a paragraph describing quality. Include one edge case.
Policies belong in the strongest layer available. A rule such as “never discuss pricing” is a request in a prompt, and a control if the pricing tool simply is not available to that agent.
ARCHITECT — context lifecycle management
Long-running agents exhaust their context. Three techniques manage this, and mature systems use all three.
Compaction. When context approaches its useful limit, summarise the conversation into a compressed state that preserves decisions, constraints, open problems and key facts, then continue from the summary. The engineering discipline is in choosing what survives: architectural decisions and unresolved issues survive; verbose tool outputs and superseded drafts do not.
Structured note-taking / externalised memory. The agent writes durable notes to a file or store outside the context window and reads them back as needed. This makes state auditable by humans and survivable across sessions — a progress file is both a memory mechanism and a handover document.
Sub-agents with clean contexts. A focused sub-agent explores a sub-problem in its own context window and returns a condensed result. The parent agent’s context stays clean; the exploration cost is isolated. This is the strongest available answer to context exhaustion in research-style work, at the price of coordination overhead and token cost (Module 8).
Just-in-time retrieval. Rather than preloading everything possibly relevant, hold lightweight identifiers — file paths, record IDs, links — and fetch content only when it is needed. This mirrors how a competent human works: you do not memorise the filing cabinet, you know where things are.
Business example
An SEO reporting agent initially received the entire crawl export, the whole analytics dump and the full brand guidelines in every run: roughly 90,000 tokens, slow, expensive and — critically — less accurate, because the genuinely important findings were competing with thousands of irrelevant rows. Re-engineered, the agent received a 400-token summary of crawl statistics plus tools to query specific issue categories on demand, and retrieved only the two brand rules relevant to reporting language. Context per run fell by roughly ninety per cent; accuracy of the priority findings improved; cost fell proportionately. Nothing about the model changed.
Common mistakes
- Adding rules to a prompt to fix behaviour, indefinitely, until the prompt is unreadable and unowned.
- Pasting entire documents “just in case”.
- Writing tool descriptions as afterthoughts.
- Treating prompt text as untracked configuration. Prompts are source code: version them, review them, test them.
- Believing a bigger context window has retired the problem.
Expert insight
The instinct to add is the enemy. When an agent misbehaves, the reflex of most teams is to add another instruction; the practice of strong teams is to ask what in the current context is producing the behaviour, and remove it. Context engineering is subtraction under constraint.
Knowledge check
- Distinguish prompt, context and agent engineering in one sentence each.
- Why is a schema better than a description of the desired output format?
- What survives compaction and what should not?
- Explain just-in-time retrieval and why it usually beats preloading.
- Why are tool descriptions considered part of the prompt?
- Give an example of moving a rule from the prompt layer to the permission layer.
Challenge
Take an existing agent or assistant configuration. Produce a context budget: list every element that enters the model’s view, with its token count. Cut thirty per cent. Measure quality on a fixed set of ten tasks before and after. Report both quality and cost.