Module 2 of 14
How Modern AI Actually Works
What is actually happening inside a large language model when it produces a response?
Learning objectives. Explain tokens, context, inference and embeddings without mathematics; describe why hallucination is structural rather than a bug to be patched; predict which failure modes come from the model and which come from your system design; and make informed decisions about model selection.
Core lesson
START HERE — the mechanism in plain English
A language model is a very large statistical system trained to predict what comes next in a sequence of text. That single sentence explains most of its behaviour, both impressive and disappointing.
Tokens. Models do not read words; they read tokens — fragments of text, roughly three-quarters of a word in English. “Unbelievable” might be three tokens. Everything is priced, limited and measured in tokens. This is why cost, latency and context limits are all really the same conversation.
Context. Everything the model can see at the moment it answers: your instructions, the conversation so far, retrieved documents, tool definitions and tool results. It is best understood as working memory — and it is the single scarcest resource in agent design. The model has no access to anything not in its context except through tools.
Inference. One pass of the model producing output. Each inference is independent. The model does not remember the last one. Every appearance of memory in every AI product you have used is engineering outside the model: previous messages re-sent as context.
Training data and cutoff. The model’s general knowledge was fixed at training time. It does not know today’s date, your prices or last week’s announcement unless something puts that information into its context.
Embeddings. A way of turning text into a list of numbers that captures meaning, so that “cancel my subscription” and “how do I stop paying” sit close together in that numeric space even though they share almost no words. This is the machinery behind semantic search and retrieval (Module 5).
Multimodality. Modern models accept and produce more than text — images, audio, documents, screenshots. Practically, this matters because much business information is not text: invoices, dashboards, product photos, whiteboards.
Reasoning. Newer models can spend additional computation “thinking” before answering, producing intermediate reasoning that improves multi-step accuracy. Treat this as a dial with a real price rather than a free upgrade: more reasoning means more tokens, more latency and more cost, and it does not help simple tasks.
PRACTITIONER — why models fail the way they do
Hallucination. The model produces a fluent, confident, wrong answer. This is structural: the system is optimised to produce plausible continuations, and a plausible continuation of “the case reference is” is a well-formed case reference. Its confidence carries no information about its accuracy. You cannot prompt this away. You can only design around it — with grounding in retrieved sources, structured output constraints, tool-based verification and human review at the points that matter.
Context rot. Model accuracy degrades as context grows, well before the technical limit is reached. Anthropic describes this as an attention budget: every token added competes with every other for the model’s limited attention, producing a performance gradient rather than a cliff. The practical rule is counter-intuitive but consistently borne out: the smallest set of high-signal tokens beats the largest set of possibly-relevant ones. More context is not better context.
Instruction dilution. A system prompt with sixty rules will not have sixty rules followed. Rules that matter belong in code, schemas or permissions, not in the twelfth paragraph of a prompt.
Non-determinism. The same input can produce different outputs. This is not a defect to be eliminated; it is the property you are buying when you choose an agent over a script. It is also why single-run testing is worthless (Module 12).
Sycophancy and anchoring. Models tend to agree with framing in their context. An agent that reads a poisoned document may adopt its framing (Module 11).
EXPERT DEEP DIVE — transformers, conceptually
The transformer’s central mechanism is attention: when processing each token, the model computes a weighted relevance over all other tokens in context. This is what gives it the ability to relate a pronoun to a noun forty sentences earlier — and also why the computation grows roughly with the square of context length, and why relevance signals get diluted as the number of competing tokens rises. Nothing in this architecture stores facts about a specific document; it stores generalised statistical structure. This is the technical reason retrieval exists.
Model selection: the practitioner’s matrix
| Consideration | Question to ask |
|---|---|
| Capability | Does the smallest model that could plausibly do this task actually pass your evaluation set? |
| Cost | What is the cost per successful task, not per token? |
| Latency | Is this interactive (seconds matter) or batch (minutes are fine)? |
| Context window | How much material must be in view at once — and can retrieval reduce it? |
| Reasoning depth | Does the task involve multi-step deduction, or is it classification and extraction? |
| Determinism needs | Would structured output plus a smaller model be more reliable than a larger model writing prose? |
| Data handling | Where does the data go, under what retention terms, in which jurisdiction? |
Route, do not standardise. Mature systems use several models: a small fast model for classification and routing, a mid-tier model for the bulk of the work, and a frontier model reserved for the genuinely hard steps. This is model routing (Module 10) and it is typically where the largest cost reductions in an agent system are found.
Business example
A firm built a document-classification agent on its most capable model, at high cost per document. Evaluation revealed that a small model with a constrained output schema achieved higher accuracy on the same test set, because the task was classification, not reasoning — and the constrained schema eliminated an entire class of formatting errors the large model was making. Cost fell by roughly an order of magnitude and accuracy rose. A better model is not the same thing as a better-fitting model.
Common mistakes
- Treating confident output as verified output.
- Filling the context window because it is available.
- Choosing the most capable model by default and never revisiting it.
- Trying to fix hallucination with stern prompt language instead of grounding and verification.
- Assuming a longer context window removes the need for retrieval. It does not; it changes the economics of the trade-off.
Expert insight
Context is a budget, not a container. The discipline of deciding what not to put in context is the highest-return skill in this whole field, and it is the direct bridge from Module 2 into Module 4.
Knowledge check
- Why does an LLM have no memory between calls, and how do products create the appearance of memory?
- Explain context rot to a non-technical colleague in two sentences.
- Why is hallucination structural rather than a fixable bug?
- What does an embedding represent, and what problem does it solve?
- Give two design responses to non-determinism.
- When would you choose a smaller model over a frontier model?
- Why is “cost per successful task” a better metric than “cost per token”?
Challenge
Take a prompt you use regularly. Cut its context by half while maintaining or improving output quality on ten test inputs. Document what you removed and what happened. This exercise converts more people to context engineering than any lecture.
Further exploration
Vaswani et al., Attention Is All You Need (2017), arXiv:1706.03762. Anthropic, Effective Context Engineering for AI Agents (2025).