Module 4 of 12
Prompting Techniques
Which techniques generalise, which are model-specific, which are overrated?
Learning objectives. Explain the major prompting techniques — zero-shot, few-shot, decomposition, classification, extraction, transformation, synthesis, comparison, critique, structured generation, iterative refinement — including what each is, why it works, when to use it, when not to, and its common failure mode; apply the research from Module 2 to separate durable technique from overrated folklore.
Why this matters
This is the module people expect to be a list of tricks. It's the opposite: it's where the research gets specific about which techniques actually earn their cost, and which are widely believed but weakly supported. Treat this module as a filter, not a menu to use in full every time.
Plain-English explanation
A "technique" here just means a pattern of asking — a way of structuring a request that reliably helps for a certain kind of task. Some patterns genuinely change outcomes. Others sound sophisticated and don't move the needle much, or actively cost you time and tokens for a gain that doesn't survive rigorous testing. Knowing which is which is the actual skill.
Core lesson — the technique catalogue
For each technique: what it is, why it works, when to use it, when not to, and its common failure mode.
Zero-shot prompting
What: asking directly, with no examples, relying on the model's general training. Why it works: modern models have broad competence from training on enormous, varied text; a clear, well-specified instruction alone is often sufficient. Use when: the task is common, well-defined, and doesn't hinge on a specific format or style only you would recognise as correct. Avoid when: the output needs to match a particular style, structure, or edge-case handling that's hard to describe but easy to show. Failure mode: vague zero-shot requests get generic, textbook-flavoured answers — the fix is almost always a clearer instruction, not switching to few-shot.
Few-shot prompting
What: including examples of the input/output pattern you want. Why it works: grounded in Brown et al.'s foundational finding that models perform new tasks from in-context examples (Module 2's Further Exploration) — but current best practice (per Anthropic's and OpenAI's own guidance) is diverse, canonical examples, not exhaustive edge-case coverage. "For an LLM, examples are worth a thousand words," as Anthropic's engineering team puts it. Use when: the desired format, tone, or edge-case handling is easier to show than describe — classification categories, a specific writing voice, a particular data-extraction shape. Avoid when: you only have one, unrepresentative example — a single bad example can anchor the model to an unwanted pattern more strongly than no example at all. Failure mode: too many similar examples narrows the model's behaviour more than intended; two or three diverse examples, including one edge case, typically beats ten similar ones.
Decomposition (task breakdown)
What: splitting a complex request into explicit sub-tasks, either in one prompt or across several turns. Why it works: reduces the chance of the model silently skipping a sub-goal buried inside a compound instruction, and makes each piece independently checkable. Use when: a request has multiple genuinely distinct parts — "analyse this data, then write a summary, then suggest three actions" is really three tasks wearing one sentence. Avoid when: the task is genuinely a single step — decomposing "rewrite this paragraph" into stages adds friction with no benefit. Failure mode: decomposing too finely turns a five-minute task into a fifteen-step ritual; decompose along genuine sub-goals, not for its own sake.
Classification
What: asking the model to sort input into predefined categories. Why it works: this is close to what models are naturally strong at — pattern-matching against learned distinctions — especially when paired with structured output (below) to guarantee a valid category comes back. Use when: you have a fixed, known set of categories (support ticket type, sentiment, lead priority). Avoid when: the categories are fuzzy, overlapping, or the true answer is genuinely "it depends" — forcing a false binary produces confidently wrong classifications. Failure mode: an ambiguous or overlapping category set gets inconsistent answers across near-identical inputs — the fix is usually clarifying the categories, not the prompt wording.
Extraction
What: pulling specific fields or facts out of unstructured text. Why it works: strongly paired with structured output (Module 6, Module 7) — a schema turns "find the invoice total" into a guaranteed-shape answer rather than a hopeful request. Use when: the source material contains the answer explicitly, even if messily formatted. Avoid when: the fact isn't actually in the source — extraction prompts will sometimes "extract" a plausible-sounding value that was never there. This is hallucination wearing an extraction costume, and it's why verification (Module 11) matters even for seemingly mechanical tasks. Failure mode: treating extraction as infallible because it "sounds mechanical." It isn't — check a sample against source.
Transformation
What: converting content from one form to another — tone, format, length, language. Why it works: this is close to the model's core training signal (predicting plausible text in a given style), so it's usually reliable with a clear target format specified. Use when: rewriting, summarising, translating, reformatting. Avoid when: the transformation requires domain judgement the model doesn't have — "simplify this legal clause" risks quietly changing its meaning, not just its wording. Always specify what must be preserved, not just what should change. Failure mode: meaning drift during aggressive simplification or summarisation — mitigated by explicitly stating what must not change.
Synthesis
What: combining multiple sources or ideas into one coherent output. Why it works: the model is genuinely good at finding structure and connections across supplied material — when that material is actually given to it (Module 6 covers grounding this properly). Use when: you have several documents, notes, or data points and need one coherent view. Avoid when: the sources conflict and you haven't told the model how to handle that — it will often silently pick one and not tell you a conflict existed. Failure mode: silent resolution of contradictory sources. Always instruct: "if sources disagree, say so explicitly rather than picking one."
Comparison
What: evaluating two or more things against shared criteria. Why it works: forcing explicit criteria (rather than "which is better?") produces a structured, checkable answer instead of a vibe. Use when: vendor comparisons, option analysis, before/after evaluation. Avoid when: you haven't defined the criteria — an undefined "which is better" invites the model to invent its own weighting, which may not match yours. Failure mode: unstated criteria produce a comparison that looks rigorous but reflects the model's implicit priorities, not yours.
Critique
What: asking the model to review and find flaws in a piece of work (yours or its own). Why it works: genuinely useful for catching surface issues — but has a real limitation covered fully in Module 11: a model reviewing its own output shares its own blind spots. External critique (a different model, a different prompt, or a human) catches more than self-critique. Use when: a second pass on drafts, code, or reasoning — treated as one useful check, not the only one. Avoid when: you're relying on it as your sole quality control for something consequential. Failure mode: confident self-approval — the critique agrees with the work because it shares the same assumptions that produced it.
Structured generation
What: constraining output to a defined shape — a schema, a fixed template, a specific set of fields. Why it works: per OpenAI's own distinction (Module 6 develops this fully): "reply in JSON" is a request; schema-constrained generation is a guarantee. This eliminates a whole class of downstream parsing failures. Use when: the output feeds into another system, a spreadsheet, or any process that needs a reliable shape. Avoid when: the task is genuinely open-ended prose where forcing structure would hurt quality (a persuasive essay doesn't want to be five rigid bullet fields). Failure mode: over-structuring creative or nuanced writing into a checklist-shaped, lifeless result.
Iterative refinement
What: treating the first response as a draft, then giving specific feedback to improve it, rather than starting over. Why it works: the model retains the context of what it produced and why, so targeted feedback ("make the second paragraph more concrete, keep the rest") is far more efficient than a fresh prompt. Use when: almost always, for anything above trivial stakes — this is one of the most underused techniques by beginners, who tend to discard and restart rather than refine. Avoid when: the first draft is so far off that starting over with a clearer prompt is genuinely faster than patching it. Failure mode: vague refinement feedback ("make it better") gets vague improvement. Be as specific about what to change as you were about the original task.
Verification (introduced here, developed fully in Module 11)
What: independently checking a claim or result before relying on it — against a source, a calculation, or a second look. Why it works: confidence is not evidence (Module 1). This is the single highest-return reliability technique available and it is not really a "prompting" technique at all — it's a habit that sits outside the prompt. Use when: anything that will be acted on, published, or relied upon. Avoid when: genuinely never — the depth of verification should scale with the stakes, not disappear. Failure mode: verifying with the same faulty assumption that produced the error in the first place. Verify against reality, not against the model's own restated confidence.
▸ ADVANCED — what the evidence actually says about the popular ones
Two findings from Module 2's research deserve restating here because this module is where they change your behaviour most directly:
Chain-of-thought ("think step by step") is not a universal upgrade. The Wharton Generative AI Labs study (June 2025) found real gains for non-reasoning models (Gemini Flash 2.0 +13.5%, Claude Sonnet 3.5 +11.7%) but negligible or even negative gains for dedicated reasoning models, at a real cost of 20–80% more time. Practical rule: use explicit "think step by step" framing for models without built-in reasoning, on genuinely multi-step tasks — skip it for simple classification/extraction, and skip it entirely on reasoning-model products where the thinking already happens internally.
Role-play framing ("You are a world-class expert...") is popular and weakly evidenced. It occasionally shifts tone or register usefully, but there is no rigorous evidence it reliably improves accuracy the way a clear objective and real context do. Treat it as a tone tool, not a capability unlock.
Business example
A customer support lead was using an elaborate multi-technique prompt (role-play persona, forced chain-of-thought, five examples) for a simple task: classifying incoming tickets into six known categories. Accuracy was fine but slow and expensive at volume. Stripped back to a direct classification prompt with a clear category list, one example per category, and a structured-output schema (no persona, no forced reasoning), accuracy held and cost per ticket dropped by roughly two-thirds. The lesson: match technique to task, not to how sophisticated the prompt looks.
Practical exercise — improve a vague prompt
Take this prompt: "Analyse our sales data and tell me what's happening." Identify which technique(s) from this module actually fit the underlying need (likely: decomposition into specific questions, plus structured output for the findings), rewrite it accordingly, and note what you added and why.
Common mistakes
- Using chain-of-thought by default on every task, regardless of whether the model or the task needs it.
- Treating self-critique as sufficient quality control for consequential output.
- Stacking multiple techniques (persona + few-shot + forced reasoning) when the task only calls for one.
- Assuming a technique that worked well on one task will transfer to a different kind of task without re-checking.
Expert insight
Every technique in this catalogue is a tool with a domain of usefulness, not a universal enhancement. The practitioners who get the most out of AI aren't the ones who know the most techniques — they're the ones who correctly diagnose which one or two the task actually needs, and stop there.
Knowledge check
- Name three techniques from this module and, for each, one situation where it would actively hurt rather than help.
- What does the research say about chain-of-thought prompting on reasoning-model products specifically?
- Why is self-critique alone insufficient as quality control for consequential output?
- What's the risk of using too many few-shot examples, or examples that are too similar to each other?
- Why does silent resolution of conflicting sources matter, and how do you prevent it?
Module summary
Prompting techniques are tools with specific domains of usefulness, not universal upgrades — the evidence is explicit that some popular techniques (chain-of-thought on reasoning models, elaborate role-play) carry real costs without reliable accuracy gains. Diagnose the task, apply the one or two techniques that actually fit, and treat verification as the technique that never turns off.
Further exploration
Meincke, Mollick, Mollick & Shapiro, The Decreasing Value of Chain of Thought in Prompting (Wharton Generative AI Labs, June 2025).