← Prompt & Context Engineering overview

Module 7 of 12

Advanced Context Architecture

How do the different layers of context actually interact?

Learning objectives. Distinguish the different layers of context — system, persistent, task, retrieved, conversational, memory, tool, environmental, user, and business-rule context; understand how these layers interact and compete for attention; and build a conceptual map of a well-architected context assembly. (Primary level: ADVANCED — this module is written for Level 3 practitioners and builders; Levels 1–2 may prefer to skim the Plain-English section and move to Module 8.)

Why this matters

Module 5 and Module 6 taught context engineering for a single request. Real systems — an AI assistant configured for ongoing use, an agent, a workflow — assemble context from multiple layers at once, each with a different lifetime and source. Understanding how these layers stack, and which one should win when they conflict, is what separates someone who writes good individual prompts from someone who can design a context system that stays reliable across hundreds of uses.

Plain-English explanation

Think of it like briefing a new employee, but broken into the different kinds of information they'd actually draw on: the standing rules of the company (system/persistent context), what they've been specifically asked to do today (task context), the specific files relevant to this task (retrieved context), what's already been discussed in this conversation (conversational context), things they've learned and retained from before (memory), what tools they have available (tool context), where they're actually operating — a live client call vs. an internal draft (environmental context), who they're talking to (user context), and the non-negotiable rules that override everything else (business-rule context). No single layer tells the whole story; the assembly of all of them does.

Core lesson

The context layers, defined

LayerContainsTypical lifetime
System / persistent instructionsIdentity, standing tone, non-negotiable rulesEvery interaction, until deliberately changed
Task contextThe current specific objective and its acceptance criteriaThis request only
Retrieved / knowledge contextDocuments, records, or data fetched for this specific questionThis request, or until the underlying source changes
Conversational contextWhat's been said earlier in this sessionThis session
MemoryFacts retained across separate sessions (preferences, past decisions)Long-term, until corrected or expired
Tool contextWhat capabilities are available and how to use themPersistent, updated when tools change
Environmental contextWhere/how this is happening — draft vs. live, internal vs. client-facingSituational
User contextWho you're serving right now — their role, history, preferencesPer-user, persistent
Business-rule contextCompliance, brand, and policy rules that override other layersPersistent, authoritative

PRACTITIONER — how the layers actually interact

The layers are not independent — they compete for the same finite attention budget described in Module 5, and they can genuinely conflict. A well-designed system resolves that with a clear priority order, stated explicitly rather than left implicit: business-rule context (compliance, non-negotiable policy) should generally override everything else; system/persistent instructions come next; task context defines what's actually being asked right now; everything else supports it.

A frequent, avoidable failure: a system prompt with dozens of accumulated rules, each added to patch a past incident, competing for the model's attention with the actual task at hand. This is the same "instruction dilution" problem Module 5 describes for documents, applied to standing rules — a system prompt with sixty rules will not have sixty rules reliably followed. Where a rule can be enforced structurally instead of just stated — a fixed dropdown instead of "please always pick from these five options," a validation step instead of "please double-check your maths" — enforce it structurally. A stated rule is a request. A structural constraint is a guarantee, and this exact distinction is the reason the companion Holistic Agent course's Security, Governance and Control module treats "policy belongs in the strongest available layer" as a governing principle — see AI agent governance — at the agent level too.

ADVANCED — long-context management techniques

For sessions or workflows that run long enough to approach real context limits, three techniques (all directly from Anthropic's engineering practice, cited fully in Module 5) manage the problem rather than ignore it:

Compaction — when context approaches a useful limit, summarise the session into a compressed state that preserves decisions, open questions, and key facts, then continue from that summary rather than the full history. The discipline is in choosing what survives: decisions and unresolved issues survive; verbose intermediate output and superseded drafts don't.

Structured note-taking — writing durable notes to a file or store outside the context window, and reading them back only when needed, rather than keeping everything live in the conversation. This is both a technical mechanism and a practical governance habit: a written progress note is auditable by a human in a way that "it's somewhere in a long chat" is not.

Just-in-time retrieval — holding lightweight references (a document title, a record ID) rather than the full content, and fetching the full content only at the moment it's actually needed. This mirrors how a competent person actually works: you don't memorise a filing cabinet, you know where things are and pull the specific file when it matters. This is the single technique most responsible for the SEO reporting business example's ninety-per-cent context reduction in Module 5.

Business example

A professional services firm's internal research assistant was configured with a system prompt that had grown, incident by incident over eight months, to include forty-one separate standing rules — some redundant, some contradicting more recent additions, none removed. Behaviour had become inconsistent and hard to predict. The fix wasn't adding a forty-second rule; it was an audit: rules were consolidated to twelve genuinely load-bearing ones, five were moved into a structural constraint (a required output template, enforced by the surrounding tool rather than requested in prose), and the rest were deleted. Consistency improved immediately — not because the AI got smarter, but because its persistent context layer stopped competing with itself.

Practical exercise — build a reusable context package

For a task you do repeatedly (client onboarding emails, weekly reports, meeting prep), map out what belongs in each layer: what's persistent (tone, standing rules), what's task-specific (this week's actual content), what's retrieved (which document or data source), and what's environmental (draft for review vs. ready to send). Write this out as a template you can reuse — this is the seed of the Context Library later in this course.

Common mistakes

  • Letting a system prompt accumulate rules indefinitely instead of periodically auditing and consolidating it.
  • Stating a rule in prose when it could be enforced structurally instead.
  • No explicit priority order when layers conflict, leaving the resolution to chance.
  • Treating conversational context as free — a very long, meandering conversation degrades the same way an overloaded single prompt does (Module 5).

Expert insight

The instinct when something goes wrong in a long-running system is almost always to add another instruction. The stronger instinct, once this module has landed, is to ask which layer the missing constraint actually belongs in — and whether it should be a rule at all, or a structural guarantee instead.

Knowledge check

  1. Name the nine context layers and give one example of each from your own work.
  2. What should generally take priority when two context layers conflict?
  3. Why is a stated rule in a system prompt described as "a request" rather than "a guarantee"?
  4. Explain compaction, structured note-taking, and just-in-time retrieval in one sentence each.
  5. Why did the professional-services firm's fix work — what actually changed, mechanically?

Module summary

Real, ongoing AI use assembles context from multiple layers with different lifetimes and sources, all competing for the same finite attention. A well-architected system states an explicit priority order, enforces what it can structurally rather than just in prose, and actively manages long-running context with compaction, note-taking, and just-in-time retrieval rather than letting it accumulate indefinitely.

Further exploration

Anthropic, Effective Context Engineering for AI Agents (2025) — the source for compaction, structured note-taking, and just-in-time retrieval as named techniques.