Module 5 of 12
Context Engineering
Why doesn't more context mean better context?
Learning objectives. Define context engineering and distinguish it precisely from prompt engineering; explain context rot and why it happens architecturally, not just anecdotally; apply the principle that more context is not automatically better context; and practise context selection, ordering, and compression on a real example.
Why this matters
This is the flagship module of the entire course. If Module 2 changes how you write a single request, this module changes how you think about everything surrounding that request — and the research behind it is some of the strongest and most consequential in the whole curriculum.
Plain-English explanation
Prompt = what you tell the AI to do. Context = what the AI needs to know and understand in order to do it well.
A prompt is one message. Context is everything the AI can actually see when it generates a response: your instructions, the conversation so far, any documents or data you've provided, examples, and (in agent settings, covered in Module 10) tool definitions and results. Context engineering is the discipline of deciding, deliberately, what belongs in that view — and, just as importantly, what doesn't.
Core lesson
PLAIN ENGLISH — context is a budget, not a container
It's tempting to think of an AI's context window the way you'd think of a filing cabinet: bigger is better, more storage is more useful, why not put everything in. The research says otherwise, and it says so for a structural reason, not just a stylistic preference.
Anthropic's engineering team, in Effective Context Engineering for AI Agents (September 2025), describes context as a finite attention budget: every token you add competes with every other token for the model's limited attention. This isn't a vague metaphor — it follows from how the underlying architecture works (the Advanced section below has the mechanism). The practical result is context rot: accuracy degrades as context grows, gradually, well before you hit any technical size limit.
This is independently confirmed by a separate, highly-cited research paper — Liu et al.'s Lost in the Middle (2023/2024) — which found that models reliably retrieve information placed at the very start or very end of their context, and reliably lose accuracy on information buried in the middle, even in models explicitly built for long context. A bigger context window is not the same claim as effective use of that window.
The rule that follows directly from both findings: the smallest set of high-signal information beats the largest set of possibly-relevant information. This is not a stylistic preference. It's what the evidence shows actually happens inside these systems.
PRACTITIONER — what to include, what to exclude, how to order it
What information enters context:
- Your actual current situation (not a generic version of it)
- Facts the AI cannot infer or shouldn't guess at
- Examples that show, not just tell, what you want
- Anything genuinely necessary for THIS response
What should be excluded:
- Background that doesn't change the answer ("just in case" material)
- Redundant restatements of something already established
- Entire documents when a relevant excerpt would do
- Old conversation history that's no longer relevant to the current ask
Structure and ordering matter, not just content. Given the "lost in the middle" finding, if you must include a large amount of material, put the single most important instruction or fact at the very start or very end of your prompt — not buried in the middle of a long document dump.
What should persist vs. what should be temporary: information that will matter across many future requests (your business's tone of voice, standing constraints) belongs in a persistent instruction layer (a system prompt or saved custom instructions) rather than being retyped every time. Information relevant only to this one request belongs only in this one request, so it doesn't quietly become "sticky" context that dilutes every future answer.
Compression, in plain terms: when you have a lot of material, don't paste all of it — summarise it to the parts that matter, or (in a Level 3 setting) let the AI fetch specific pieces on demand rather than front-loading everything (Module 7 develops this as "just-in-time retrieval").
ADVANCED — why this happens architecturally
The transformer architecture underlying modern language models computes relevance across every pair of tokens in context — a computation that grows roughly with the square of context length, and where every additional token is, in principle, competing for relevance against every other. Most training data historically consisted of shorter sequences than today's advertised context windows, meaning models have seen proportionally less training signal on how to use very long contexts well, compared to how much they've seen on shorter ones. This is the structural reason context rot is a gradient, not a cliff — degradation begins well before you hit any technical limit, and it degrades gradually rather than failing suddenly.
Practical implication for anyone building repeatable systems: measure quality against context size directly, on your own tasks, rather than assuming "the window is big enough" settles the question. Module 4's own Advanced example (support-ticket classification) is exactly this in practice — stripping unnecessary context and technique both improved accuracy and cut cost.
Beginner example
Context-heavy and unfocused: pasting an entire 20-page company handbook and asking "how should I respond to this customer complaint?"
Context-engineered: extracting the two relevant policy paragraphs (refunds and response-time SLA) and pasting just those, with the actual complaint, and the instruction "respond per our refund policy above; keep it under 150 words."
The second version isn't just shorter to write — per the research above, it's genuinely more likely to produce an accurate, on-policy answer, because the relevant information isn't competing with eighteen irrelevant pages for the model's attention.
Intermediate example
A recruiter drafting candidate outreach used to paste an applicant's entire CV plus the full job description into every message-drafting prompt. Rebuilt: a two-line context block ("Candidate: 6 years in B2B sales, currently at [Company], strongest recent achievement: exceeded quota by 40%. Role: Senior AE, remote-first, our differentiator is a 4-day work week.") outperformed the full-document version on every draft reviewed, because the message stopped drowning in irrelevant detail and started foregrounding the two facts that actually mattered for a compelling first message.
Advanced example
A technical team building an internal research assistant initially preloaded every user's query with the full text of the top ten potentially-relevant internal documents "to be safe." Accuracy on specific questions was worse than expected, and costs were high. Following Anthropic's just-in-time retrieval pattern (fully developed in Module 7), they switched to giving the model lightweight references — document titles and short summaries — with a tool to fetch the full text of a specific document only when needed. Accuracy improved and token cost fell by an order of magnitude, because the model was no longer required to find a needle inside ten haystacks it didn't need in the first place.
Business example
An SEO reporting workflow (a real, documented pattern) originally sent an AI the entire monthly crawl export and full analytics dump for every client report — around 90,000 tokens per run — reasoning that more data meant a more thorough report. Reports were slow, expensive, and — the counterintuitive finding — less accurate on the findings that mattered, because genuinely important issues were competing with thousands of irrelevant data rows for attention. Re-engineered to send a compact summary of crawl statistics plus the ability to query specific issue categories on demand, context per run fell by roughly ninety per cent, accuracy on priority findings improved, and cost fell proportionately. Nothing about the underlying AI model changed — only what it was shown.
Holistic Agent Application
Inside Holistic Agent, every agent you configure is trained on the documents you upload as its knowledge base. This module's principle applies directly: uploading your entire company handbook, every old policy version, and unrelated internal documents to a single customer-facing agent doesn't make it smarter — per the research above, it's more likely to dilute the agent's attention on the documents that actually answer customer questions. When you set up an agent's knowledge base (Module 9 walks through this step by step), curate what you upload the same way you'd curate a prompt: the smallest set of genuinely relevant, current documents, not an exhaustive archive "just in case."
Practical exercise — remove irrelevant context
Take a prompt you regularly send with a large amount of pasted background (an email thread, a long document, extensive notes). Cut the pasted material by at least half, keeping only what's genuinely load-bearing for the specific question you're asking. Run both versions on the same request and compare the results — not just for speed, but for whether the more focused version is actually as good or better.
Common mistakes
- Assuming a bigger context window removes the need to curate what you put in it — it changes the economics, not the underlying attention-competition problem.
- Pasting entire documents "just in case" instead of relevant excerpts.
- Letting old, no-longer-relevant conversation history quietly accumulate in a long chat instead of starting fresh when the topic genuinely changes.
- Burying the most important instruction in the middle of a long prompt instead of the start or end.
Expert insight
When an AI's output disappoints, the instinct of most people is to add more instructions, more caveats, more background — to over-correct with volume. The stronger instinct, once you've internalised this module, is the opposite: ask what in the current context is producing the unwanted behaviour, and remove it. Context engineering is subtraction under constraint, not accumulation.
Knowledge check
- State the difference between a prompt and context in one sentence each.
- What is context rot, and what architectural fact explains why it happens?
- Summarise the "lost in the middle" finding and its practical implication for where you place your most important instruction.
- Give an example of information that should persist across requests versus information that should stay temporary.
- Why did removing context improve accuracy in the SEO reporting business example, rather than just making it faster?
Module summary
Context engineering is the discipline of deciding what an AI needs to see — and what it shouldn't — to do a specific task well. The research is unambiguous: context is a finite attention budget, not an unlimited container, and more context reliably produces worse results past a point that arrives earlier than most people expect. The smallest set of high-signal information beats the largest set of possibly-relevant information, in evidence, not just in principle.
Further exploration
Anthropic, Effective Context Engineering for AI Agents (2025). Liu et al., Lost in the Middle: How Language Models Use Long Contexts, arXiv:2307.03172.