Module 12 of 12
From Prompt Engineer to AI Systems Thinker
What's the actual progression, and what comes next?
Learning objectives. Trace the progression from AI user to AI systems thinker; apply the evidence ladder (established / emerging / experimental / speculative) to new claims you'll encounter after this course; and know exactly what to read next, and in what order.
Why this matters
This is the module that ties off the course's central promise: that today's prompt engineering naturally develops into tomorrow's context engineering, agent instruction design, and eventually systems thinking — not because the skills change, but because the same underlying judgement gets applied at greater scale and higher stakes.
Plain-English explanation
The progression this course has walked you through, one module at a time:
"I ask AI questions." (Module 1) ↓ "I know how to instruct AI clearly." (Modules 2–4) ↓ "I know how to give AI the right context — and to leave out the wrong context." (Modules 5–7) ↓ "I can apply this systematically, across platforms and inside a real product." (Modules 8–9) ↓ "I understand what changes when instructions become persistent, autonomous agent work." (Module 10) ↓ "I can tell the difference between a result that worked and a system I can trust." (Module 11)
That last line is the actual destination. Not "I found a good prompt" — "I can engineer the environment in which an AI reliably does useful work."
Core lesson
PLAIN ENGLISH — the evidence ladder, for everything you read next
You will keep encountering confident claims about AI after this course ends — new techniques, new products, new predictions. Apply the same four-level ladder this course has used throughout:
| Label | Meaning | What to look for before believing it |
|---|---|---|
| Established | Demonstrated reliably, across sources, with real evidence | Reproducible findings, ideally from more than one independent source |
| Emerging | Working in real use, but still forming, evidence still accumulating | Credible early implementations, honestly-reported failures alongside successes |
| Experimental | Shown in research or limited settings, not yet dependable in normal use | A specific study or demo, with stated limitations |
| Speculative | A plausible story, not yet evidence | An argument, not a result — treat as a hypothesis, not a plan |
Applied to this course's own content, honestly: the core research in Modules 2, 4, and 5 (Wharton's chain-of-thought and prompt-wording findings, Anthropic's context engineering practice, the "Lost in the Middle" paper) is Established. Automatic prompt optimisation tools like DSPy (mentioned briefly in Module 4's advanced material) are Emerging — real and growing, not yet a beginner-relevant default. Claims about AI systems that fully self-improve their own prompts without human oversight are Speculative — plausible-sounding, not evidenced.
PRACTITIONER — what to actually do with a new claim
When you hit a confident new claim after this course: ask what evidence actually backs it, whether it's been tested more than once, and — the single most useful question from this entire course, applied one more time — which HPC-10 element does this claim actually change? A genuinely new technique usually sharpens one specific element (a better way to structure Context, a smarter way to do Verification). A vague claim that "this changes everything" without naming what, specifically, tends to fail that test.
ADVANCED — what's likely to remain true regardless of what comes next
Consistent with the Holistic Agent course's own The Future of Holistic Agents module, and directly supported by this course's research:
- Context will remain the binding constraint. Bigger windows change the economics, not the underlying need to curate what matters (Module 5).
- Verification will remain necessary. Confidence has never been evidence, and nothing in the research suggests that will change (Module 11).
- Clarity of objective will remain more valuable than clever wording. The Wharton findings on this are unlikely to reverse just because models improve (Module 2).
- The transition from prompt to context to agent instruction will keep being a matter of degree, not of kind. The same ten elements, under more or less rigour (Module 10).
Business example
A learner who completed this course as a marketing manager reported, six months later, that the single habit that had actually compounded wasn't a specific prompt template — it was the reflex of stating the objective before the task, and building a five-question evaluation set before trusting any new AI-assisted process at work. Neither of those requires the latest model or platform. That's the point.
Practical exercise — the final capstone
Individual track — Build Your AI Work System. Produce, for your own real work: (1) a stated objective, (2) standing instructions you'll reuse, (3) a simple context architecture (what's persistent, what's per-task), (4) your actual knowledge sources, (5) a small prompt library (drawn from this course's Prompt Library), (6) a lightweight workflow, (7) evaluation criteria, (8) a verification habit matched to your real stakes, (9) your AI-engine strategy (Module 8), and (10) — optionally — a first sketch of what an agent version might look like (Module 10).
Business track — Design an AI-Powered Business Workflow. Identify one real process in your organisation. Determine, explicitly: what should remain human, what can be automated outright, what genuinely benefits from AI judgement, what might eventually warrant a full agent, what context that would require, what safeguards it would need, and how you'd measure whether it actually worked.
Common mistakes
- Treating this course's end as "finished" rather than as the point where the same skills start compounding in a different setting.
- Chasing every new technique that appears with confident marketing language, without applying the evidence ladder.
- Building an elaborate agent before the evaluation habit from Module 11 is genuinely a habit, not a one-time exercise.
Expert insight
Models will change — meaningfully, and probably faster than most predictions suggest. The ten elements of HPC-10 won't. Purpose, instructions, and constraints will always need stating. Context, knowledge, and resources will always need curating, not maximising. Output, evaluation, verification, and iteration will always be what separates a lucky result from a trustworthy one. That's not a hedge against this course going out of date — it's the actual reason it was built around evidence and structure rather than around whichever model happens to be newest this month.
Knowledge check
- Trace the six-step progression this course has walked through, from Module 1 to Module 11, in your own words.
- Apply the evidence ladder to one AI claim you've encountered recently, outside this course.
- What question should you ask about a new AI technique before adopting it, per this module?
- Name the four things likely to remain true regardless of future model improvements, and why each is grounded in this course's research rather than just asserted.
- What comes next, specifically, if you want to continue past this course?
Module summary
This course's real destination was never "write better prompts" — it was the shift from asking questions to engineering the environment an AI operates in: its objective, its context, its knowledge, its constraints, and how you check its work. That shift, tested against real evidence throughout, is what continues to compound as models change, as your own use grows more serious, and — for those who want it — into the Holistic Agent course's full treatment of building and governing AI agents.
Further exploration
The Holistic Agent course — the direct continuation, from The New Age of AI onward, of everything this course has built toward.
REFERENCE — The Prompt Library
A small number of genuinely reusable templates beats a large number of shallow ones. Each entry states its purpose, how it works, what to customise, when to use it, and its common failure mode — apply the Prompt–Context Cycle from the introduction to adapt any of these to your own work rather than using them unedited.
1. Research brief
Purpose: gather and structure information on a topic you need to understand before deciding something. Template: "I need to understand [topic] well enough to [decision it will inform]. Research and summarise: [2–4 specific sub-questions]. Flag anything where sources disagree, rather than picking one silently. Cite where each claim came from." Customise: the sub-questions — vague research briefs produce vague research. Use when: you have a real decision riding on the answer, not idle curiosity. Failure mode: omitting the decision the research will inform — the AI can't judge relevance without knowing what "relevant" means for you (Module 2).
2. Structured analysis
Purpose: turn raw data or notes into a clear, checkable analysis. Template: "Analyse [data/notes] to answer: [specific question]. Structure your answer as: (1) headline finding, (2) supporting evidence, (3) what this doesn't tell us / open questions. Do not speculate beyond what the data supports." Customise: the specific question — "analyse this and tell me what's happening" (Module 4's worked example) is the failure mode to avoid. Use when: you need a defensible, checkable conclusion, not just a narrative. Failure mode: skipping the "what this doesn't tell us" section, which is often where the most honest and useful insight lives.
3. Summarisation with fidelity check
Purpose: condense a document without losing what matters. Template: "Summarise [document] in [length]. Preserve: [specific things that must not be lost — figures, names, caveats]. Flag anything you had to omit that seems important." Customise: the "preserve" list — this is what prevents meaning drift (Module 6). Use when: the source is long enough that summarising is genuinely useful, and accuracy of the summary matters. Failure mode: no fidelity check — a fluent, well-written summary can still silently drop something important.
4. Structured comparison
Purpose: compare options against explicit, shared criteria. Template: "Compare [A], [B], [C] against these specific criteria: [list]. Present as a table. For any criterion where information is genuinely missing, say so rather than guessing." Customise: the criteria — an undefined "which is better" invites the model's own implicit weighting (Module 4). Use when: a real decision depends on the comparison. Failure mode: letting the AI choose its own criteria — you'll get a comparison, not necessarily the one that matters to you.
5. Planning / decomposition
Purpose: break a complex goal into a concrete, ordered plan. Template: "Goal: [outcome]. Break this into ordered steps, noting dependencies (what must happen before what). Flag the step most likely to go wrong and why." Customise: the goal, stated as an outcome (Module 2) — not as the first step. Use when: a task has genuinely multiple, interdependent parts (Module 4's decomposition technique). Failure mode: decomposing a genuinely single-step task, adding friction with no benefit.
6. Strategy options with trade-offs
Purpose: generate genuinely distinct strategic options, not variations on one idea. Template: "Given [situation and constraint], propose three genuinely different approaches to [objective]. For each: the core idea, the main risk, and what would have to be true for it to work." Customise: the constraint — without one, options tend to converge on generic advice. Use when: you want to see the option space, not just one recommendation. Failure mode: accepting three options that are really the same idea with different wording — push back if that happens.
7. Writing with a defined voice
Purpose: produce on-brand, consistent writing. Template: "Write [content type] for [audience]. Objective: [what it needs to achieve]. Tone: [specific description, ideally with one example of the voice]. Length: [limit]. Avoid: [specific things to exclude]." Customise: the tone description — "professional" means little; "warm but not salesy, like a knowledgeable friend, not a brochure" means something. Use when: any writing that represents you or your brand. Failure mode: vague tone instructions produce generic, forgettable copy (Module 2's marketing plan example).
8. Editing with preserved intent
Purpose: improve a draft without changing what it's actually saying. Template: "Edit this for [specific goal — clarity, brevity, tone]. Do not change: [facts, structure, or claims that must stay as written]. Show me what changed and why, briefly." Customise: what must not change — this is the fidelity guard from template 3, applied to editing. Use when: refining your own or a colleague's draft. Failure mode: an open-ended "improve this" invites unwanted rewrites of things that were fine as they were.
9. Extraction with a schema
Purpose: pull specific fields from unstructured text reliably. Template: "From [document], extract: [field 1], [field 2], [field 3]. If a field isn't present, say 'not found' — do not guess or infer a plausible value." Customise: the field list, and always include the explicit "do not guess" instruction (Module 4 and Module 6). Use when: the output feeds into a spreadsheet, form, or another system. Failure mode: omitting "do not guess" — extraction prompts will otherwise sometimes invent a plausible-sounding value.
10. Classification with a fixed category set
Purpose: sort items into known categories consistently. Template: "Classify [item] into exactly one of these categories: [list, with a one-line definition of each]. If none fit well, say 'unclear' rather than forcing a fit." Customise: the category definitions — ambiguous categories produce inconsistent classification (Module 4). Use when: you have a genuinely fixed, known category set. Failure mode: forcing a choice when "unclear" should be a valid answer — this is where judgement gets measured (Module 11).
11. Document/audit analysis
Purpose: review a document (contract, report, policy) for specific issues. Template: "Review [document] specifically for [issue type — e.g. non-standard terms, compliance gaps, inconsistencies]. For each finding: quote the exact text, explain the concern in plain language, and rate severity (low/medium/high). Do not comment on anything outside this scope." Customise: the issue type — an unscoped "review this" produces an unfocused, less useful pass. Use when: any structured document review, always with human sign-off on anything consequential (Module 11). Failure mode: treating the output as a final judgement rather than a first pass requiring the verification hierarchy (Module 11).
12. Customer service response
Purpose: draft a response that's grounded in real policy, not generic reassurance. Template: "Using [our actual policy, pasted or referenced], respond to this customer message: [message]. Tone: [specific]. If the situation isn't covered by the policy, say so rather than improvising a commitment." Customise: the policy reference — this is Module 6's grounding discipline, applied directly. Use when: any customer-facing draft. Failure mode: letting the AI improvise a policy position it isn't actually authorised to make.
13. Project status / meeting notes
Purpose: turn a transcript or rough notes into clear, actionable output. Template: "From [transcript/notes], produce: (1) key decisions made, (2) action items with owners if stated, (3) open questions. If an owner wasn't stated for an action item, flag it rather than assigning one." Customise: rarely needs much — this template travels well as-is. Use when: after any meeting or working session worth capturing. Failure mode: letting the AI invent an owner for an action item nobody actually claimed.
14. Agent instruction starter (bridge to Module 10)
Purpose: draft the first version of a persistent agent instruction, ready for human refinement. Template: "Draft a persistent instruction for an AI agent whose mission is: [outcome, not activity]. It should handle: [in scope]. It must never: [boundaries]. When uncertain, it should: [escalation behaviour]. Its tone should be: [voice]." Customise: everything — this is a starting scaffold, not a finished instruction; review against Module 10's full checklist before using it for real. Use when: you're ready to move from a manual prompt to something more persistent (Module 10). Failure mode: treating this draft as complete without the human review Module 10 and 11 both insist on.
REFERENCE — The Context Library
Reusable context packages, not single-use pastes — build these once per subject, update them when they go stale, and reuse them across many prompts. This is the direct product of Module 7's exercise.
Business context template: "[Business name] is a [size/type] business in [industry/location]. We serve [audience]. Our current priority is [this quarter's focus]. Our tone is [voice description]."
Customer context template: "This customer: [tier/segment], history: [relevant prior interactions], currently: [their situation right now]."
Project context template: "Project: [name]. Objective: [outcome]. Current phase: [status]. Key constraint: [budget/timeline/other]. Stakeholders: [who cares about what]."
Research context template: "Research area: [topic]. What we already know: [brief]. What we're trying to find out: [specific gap]. Sources we trust most: [if relevant]."
Product context template: "Product: [name]. What it does: [one line]. Who it's for: [audience]. Key differentiator: [what makes it different, specifically, not just 'better']."
Brand context template: "Voice: [description with an example]. Never: [specific words/claims/tone to avoid]. Always: [non-negotiables]."
Company knowledge template: "Current policy on [topic]: [the actual policy, pasted, dated]. Supersedes: [what it replaced, if relevant, so old material isn't accidentally used]."
Agent knowledge template (bridge to Module 9/10): "This agent's knowledge base should include: [current, specific documents]. Should explicitly exclude: [superseded/irrelevant material]. Authority order if sources conflict: [which wins]."
Task context template: "Right now, specifically: [the actual situation this request concerns] — not the general case."
Building your own: per Module 7, a good context package separates what's persistent (rarely changes — brand voice, standing policy) from what's per-task (this week's actual figures, this specific customer). Keep them as separate, dated documents you paste in combination, not one giant merged file that goes stale unevenly.
REFERENCE — Prompt & Context Anti-Patterns
What not to do — this section carries as much weight as the techniques it warns against.
1. Vague instructions. "Write something about X" instead of a stated objective and task (Module 2). Fix: state the objective as its own sentence before the task.
2. Conflicting instructions. Constraints that quietly contradict each other ("be extremely thorough" and "keep it to two sentences"). Fix: read your own prompt as if you were the AI — would you know which instruction wins?
3. Unnecessary verbosity. A five-paragraph prompt for a task that needed one sentence (Module 3's over-engineering trap). Fix: match structure to stakes and repeatability.
4. Irrelevant context. Pasting an entire document when an excerpt would do (Module 5). Fix: the smallest set of high-signal information, always.
5. Missing context. Assuming the AI knows your situation because you've mentioned it before, in a different conversation (Module 6). Fix: if it's not in front of the AI right now, it doesn't know it.
6. Too many tasks at once. "Analyse this, then write a summary, then draft three social posts, then suggest a hashtag strategy" in one breath. Fix: decompose explicitly (Module 4), or run as separate requests.
7. Excessive role-playing. Elaborate personas ("You are a world-renowned expert with 30 years of experience...") in place of a clear objective and real context. Fix: per Module 4's research note, role-play shifts tone, not reliably accuracy — spend your effort on the objective instead.
8. Unsupported assumptions. Assuming the AI will infer your unstated priorities, your industry's norms, or your risk tolerance. Fix: state it; don't assume it's obvious.
9. No evaluation. Trusting a prompt because it worked once (Module 11). Fix: a small test set, always, for anything reused.
10. No verification. Treating confident output as checked output (Module 1, Module 11). Fix: match verification depth to stakes, every time.
11. Blindly copying prompts. Using a template from this library, or anywhere else, unedited for your actual situation. Fix: every template here says "customise" for a reason — treat that as a requirement, not a suggestion.
12. Overengineering simple tasks. Building an elaborate context package and multi-step technique stack for "rewrite this sentence to be friendlier." Fix: Module 3's judgement call — not everything needs the full structure.
13. Using an agent where a simple, reliable process would do. Building autonomous complexity for a task that's actually the same five steps every time. Fix: the Holistic Agent course's Agency Test — is the path genuinely unpredictable, does it need real judgement, can success be checked? If not, a straightforward process beats an agent.
REFERENCE — The Prompt & Context Optimisation Loop
Introduced in the course overview as the Prompt–Context Cycle; here is the full diagnostic version, including how to tell where a failure actually came from.
DESIGN → TEST → OBSERVE → DIAGNOSE → MODIFY → RETEST → EVALUATE → STANDARDISE
When something goes wrong, diagnose the actual source before you fix anything — this table, an expansion of HPC-10's diagnostic table from the introduction, sorts failures by where they actually originate:
| Where the failure might be | How to tell | What to check first |
|---|---|---|
| The prompt itself | The instruction is genuinely ambiguous, even to a careful human reading it cold | Module 2–3: is the objective stated? Is the task specific? |
| The context | The AI answered a generic version of your situation | Module 5–6: was your actual situation, and the relevant material, actually supplied? |
| The model | The same well-structured prompt performs differently across engines | Module 8: is this a task where model choice genuinely matters? |
| Missing information | The AI needed a fact it was never given and had to guess | Module 6: was the knowledge source actually complete for this question? |
| A tool or capability gap (agent settings) | The AI "wanted" to do something it had no way to actually do | Module 10: does its resource/tool list match what the task needs? |
| Genuine ambiguity in the task itself | Even you're not sure what a "good" answer would look like | Module 2: the objective isn't actually clear yet — clarify it before blaming the AI |
| The evaluation criteria | Different people judge the same output differently | Module 11: your success criteria were never made explicit |
THE CAPSTONE
See Module 12 for the full brief — Build Your AI Work System (individual track) and Design an AI-Powered Business Workflow (business track). Both are scored against a completed HPC-10 canvas: every one of the ten elements should be identifiable, specifically, in your submission, not just implied.
CERTIFICATION FRAMEWORK
Five levels, defined by demonstrated competence, deliberately numbered to hand off cleanly into the Holistic Agent course's own five-level path.
Level 1 — AI Interaction Fundamentals
Modules: 1–3. Competency: explains what an AI model is and isn't; distinguishes question, prompt, instruction, workflow, and agent; writes a well-formed prompt using Goal-Context-Task-Constraints-Inputs-Output-Evaluation, matched appropriately to the task's stakes. Assessment: Knowledge checks 1–3; the Module 3 exercise, showing both a minimal and a fully-structured prompt with a stated reason for the difference.
Level 2 — Prompt Engineering Practitioner
Modules: 1–4. Competency: selects the right technique(s) for a task from the Module 4 catalogue; explains, with evidence, why chain-of-thought and role-play are not universal upgrades; builds and reuses a small personal prompt library. Assessment: the Module 4 exercise; three entries from the learner's own Prompt Library, each with a stated purpose and failure mode.
Level 3 — Context Engineering Practitioner
Modules: 1–7. Competency: applies context rot and the "smallest high-signal set" principle in practice; grounds AI output in real documents with explicit conflict-handling; builds a multi-layer context package (Module 7) for a real, repeated task. Assessment: the Module 5 and Module 7 exercises; a demonstrated before/after context reduction with measured quality difference.
Level 4 — Advanced AI Interaction Designer
Modules: 1–9. Competency: routes tasks across AI engines deliberately; configures and evaluates a real HolisticAgent.ai agent using the full HPC-10 canvas; separates model-agnostic design from engine-specific tuning. Assessment: the Module 8 and Module 9 exercises; a completed HPC-10 canvas for one live agent configuration, including a ten-question evaluation run.
Level 5 — Agent Instruction & Context Architect
Modules: all twelve, plus the capstone. Competency: drafts a genuine agent instruction set distinguishing it clearly from a single prompt; builds and maintains an evaluation set including refusal cases; explains prompt injection and its architectural (not wording-level) defence; applies the evidence ladder to new claims independently. Assessment: the full capstone (either track); the Module 10 exercise; a completed evaluation set (Module 11) for a real, repeated task, including at least two refusal cases. This level is the explicit, intended handoff into the Holistic Agent course's own Level 1 — AI & Agent Foundations.
GLOSSARY
Plain English first; technical precision second. Terms already defined in the Holistic Agent course's own glossary are marked (shared) and kept consistent with that definition rather than restated differently.
AI (Artificial Intelligence). (shared) Software that performs tasks normally requiring human intelligence.
Agent. (shared) A system that interprets an objective, decides its own steps, uses tools, takes action, and observes the results — covered at the introductory level in Module 10; developed fully in the Holistic Agent course's Anatomy of an AI Agent module.
Chain of thought. Producing intermediate reasoning before answering. Genuinely helpful for non-reasoning models on multi-step tasks; shown by research to add cost with little or no accuracy gain on dedicated reasoning models (Module 4).
Compaction. Summarising a long session into a compressed state that preserves decisions and open problems, then continuing from it (Module 7).
Context. (shared) Everything an AI can see at the moment it answers — the scarcest resource in effective AI use (Module 5).
Context engineering. Deciding what belongs in an AI's view at each step, and — more importantly — what doesn't (Module 5).
Context rot. Degradation of accuracy as context grows, well before any technical size limit is reached (Module 5).
Decomposition. Breaking a complex task into explicit, ordered sub-tasks (Module 4).
Evaluation. (shared) Systematic measurement of an AI's output against known-good examples, rather than trusting a single result (Module 11).
Few-shot prompting. Including examples of the input/output pattern you want, to guide behaviour without retraining (Module 4).
Grounding. (shared) Basing AI output on real, supplied material rather than generalised training — the primary defence against hallucination (Module 6).
Hallucination. (shared) Fluent, confident, incorrect output. Structural to how these systems work, not a bug to be prompted away (Module 1).
HPC-10. The Holistic Prompt & Context Canvas — this course's ten-element framework (Purpose, Instructions, Constraints, Context, Knowledge, Resources, Output, Evaluation, Verification, Iteration), organised into three layers (Intent, Information, Assurance). The deliberate, smaller sibling of the Holistic Agent course's HAC-20.
In-context learning. A model's ability to perform a new task from examples or instructions given in the prompt itself, without retraining — the mechanism behind few-shot prompting (Module 4).
Inference. (shared) One pass of a model producing output; stateless — models don't remember between separate calls (Module 1).
Iteration. In HPC-10: capturing what was learned from a result and applying it to the next attempt, rather than re-deriving from scratch each time.
Just-in-time retrieval. Holding lightweight references and fetching full content only when actually needed, rather than pre-loading everything (Module 7).
LLM (Large Language Model). (shared) A model trained on large amounts of text to predict plausible continuations; the technology behind ChatGPT, Claude, and Gemini (Module 1).
LLM-as-judge. Using a model to score another output against a rubric — scales evaluation, but requires human-calibrated rubrics and carries documented biases (Module 11 references this; developed fully in the Holistic Agent course's Evaluating AI Agents module).
Multimodal. Able to process more than text — images, PDFs, audio (Module 1).
Prompt. (shared) A deliberate request specifying a task, as distinct from a casual question or a persistent instruction (Module 1).
Prompt engineering. (shared) The deliberate design of instructions and inputs that guide an AI system toward a desired outcome (Module 2).
Prompt injection. (shared) Malicious or misleading instructions hidden in content an AI reads, attempting to redirect it from the operator's actual intent. No complete wording-level fix exists (Module 11).
RAG (Retrieval-Augmented Generation). (shared) Fetching relevant material and placing it in context so an AI answers from real sources rather than generalised training (Module 6).
Reasoning model. A model that spends extra computation "thinking" before answering — a genuine capability with a real cost, not a free upgrade (Module 1).
Structured output. (shared) Constraining generation to a defined schema, converting a request for a format into a guarantee of one (Module 4, Module 6).
System prompt / persistent instructions. Standing instructions that apply to every interaction with an AI assistant or agent, until deliberately changed (Module 7).
Token. (shared) A fragment of text — roughly three-quarters of a word in English — the unit context, cost, and limits are measured in (Module 1).
Verification. (shared) Independently checking a claim or result before relying on it — the highest-return reliability practice available, and distinct from asking the AI to check itself (Module 4, Module 11).
Zero-shot prompting. Asking directly, with no examples, relying on the model's general training (Module 4).
RESEARCH REFERENCES
Primary and near-primary sources used in this course's research phase. Time-sensitive claims (protocol details, live pricing, platform features) should always be checked against their current primary source before being relied upon.
Research papers
- Brown et al., Language Models are Few-Shot Learners, arXiv:2005.14165 (2020) — the foundational in-context learning result.
- Liu et al., Lost in the Middle: How Language Models Use Long Contexts, arXiv:2307.03172 (2023; TACL 2024) — positional degradation in long-context use.
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv:2005.11401 (2020) — the origin of RAG.
- Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, arXiv:2201.11903 (2022) — the original CoT result, read alongside the Wharton report below for its current limits.
Applied research
- Meincke, Mollick, Mollick & Shapiro (Wharton Generative AI Labs), Prompt Engineering is Complicated and Contingent (4 March 2025) — rigorous, repeated-trial evidence on what prompt wording actually does and doesn't reliably change.
- Meincke, Mollick, Mollick & Shapiro (Wharton Generative AI Labs), The Decreasing Value of Chain of Thought in Prompting (8 June 2025) — CoT's shrinking benefit on reasoning models specifically.
Engineering practice
- Anthropic, Effective Context Engineering for AI Agents (29 September 2025) — the primary source for context rot, compaction, structured note-taking, and just-in-time retrieval.
- Anthropic, Building Effective Agents — workflow-versus-agent distinction, referenced in Module 10's handoff to the Holistic Agent course.
- OpenAI, official Prompt Engineering developer guidance — the Identity→Instructions→Examples→Context structure and model-specific guidance referenced in Modules 3, 4, and 8.
- OpenAI, Introducing Structured Outputs in the API — the request-vs-guarantee distinction underlying Module 6's structured-generation content.
Security
- OWASP, Top 10 for LLM Applications (2025 edition) — the current standard reference for prompt injection, cited in Module 11.
Benchmarks, protocol specifications, and platform features referenced throughout (particularly Module 9) should be verified against their live, current source — this course reflects a snapshot as of 31 August 2026.
FURTHER LEARNING
Finished this course? The direct next step is the Holistic Agent course, published in full — start at The New Age of AI. You already have the vocabulary; what's genuinely new there is autonomy, tools that act on the real world, multi-agent coordination, and governance at the system level, all built on the same evidence-first, three-depth teaching approach this course has used throughout.
For hands-on practice ahead of that: configure and evaluate a second live HolisticAgent.ai agent (Module 9's exercise, repeated on a different agent type), and keep growing your personal Prompt Library and Context Library — they are the most durable artefacts this course produces, and they'll keep paying off long after any specific model you're using today has been replaced.
Prompt & Context Engineering — Edition 1.0, August 2026. HolisticAgent.ai / Holistic Agent Academy.
This edition reflects the state of research and the HolisticAgent.ai platform at the time of writing. Model capabilities, platform features, and live pricing change — verify time-sensitive claims, particularly anything in Module 9, against the current live platform before relying on them. The evidence, frameworks, and design discipline in this course are intended to outlast any single model release.