← Course overview

Module 5 of 14

Agent Memory and Knowledge

How does an agent remember, and what should it be allowed to forget?

Learning objectives. Distinguish the memory types and choose the right one per use case; explain retrieval-augmented generation and its failure modes; decide what an agent should remember and what it must forget; and design memory governance that satisfies privacy obligations.

Core lesson

Two things are constantly confused. Knowledge is what the organisation knows — documents, records, policies, data. It exists whether or not the agent runs. Memory is what the agent retains from its own experience — decisions, preferences, state, history. Knowledge is a library; memory is a notebook. They have different owners, different lifecycles, different risks, and different regulatory treatment.

START HERE — the memory types, by analogy

TypeHuman analogyIn an agentTypical lifetime
Working memoryWhat you are holding in mind right nowThe context windowOne step
Short-term memoryThis conversationRecent turns, current task stateOne session
Episodic memory“Last Tuesday we agreed X”Logged past runs and their outcomesWeeks–months
Semantic memoryFacts you know about the worldStored facts about clients, products, entitiesLong
Procedural memoryKnowing how to ride a bikeLearned or curated workflows, playbooks, tool-use patternsLong

PRACTITIONER — retrieval and RAG

Retrieval-augmented generation means fetching relevant material and placing it in context so the model answers from sources rather than from generalised training. It is the standard answer to two problems: knowledge the model never had, and knowledge that changes.

The mechanics: documents are split into chunks; each chunk is converted to an embedding and stored in a vector database; a query is embedded; the nearest chunks are retrieved and inserted into context, ideally with citations.

RAG fails in specific, diagnosable ways, and the failures are almost never in the model:

FailureCauseResponse
Right answer not retrievedPoor chunking, poor embeddings, vocabulary mismatchHybrid keyword + semantic search; better chunk boundaries; query rewriting
Retrieved but ignoredToo much competing contextRetrieve fewer, better chunks; rerank
Retrieved but staleNo freshness signal or reindexingTimestamp chunks; expire and reindex
Confidently wrong citationNo verification of quote against sourceRequire verbatim quotes; check them programmatically
Contradictory sourcesNo authority hierarchyRank sources; state the conflict rather than resolving it silently

Hybrid search matters more than most teams expect. Semantic search finds meaning but misses exact identifiers — product codes, invoice numbers, error strings. Keyword search finds those and misses paraphrase. Production systems run both and merge.

Structured beats unstructured whenever it exists. If the answer is a row in a database, query the database. RAG over a PDF export of a table is a common and entirely avoidable self-inflicted wound.

Knowledge graphs store entities and the relationships between them. They are the right tool when the question is about connections — which clients are affected by this supplier, which policies govern this contract — rather than about passages of text. They cost more to build and repay it in multi-hop reasoning.

ARCHITECT — memory design and governance

Three questions decide every memory design:

  1. What is worth remembering? A fact is worth persisting if it will change a future decision. Preferences, decisions, entity facts and outcomes qualify. Chat transcripts, verbose tool output and superseded drafts do not.
  2. How is it corrected? Memory that cannot be corrected becomes an accumulating error surface. Every memory store needs an inspection view, an edit path and a named owner.
  3. How does it expire? Design forgetting deliberately: time-to-live per class of memory, supersession rules, and hard deletion on request.

Memory is a security boundary. Memory poisoning — an attacker planting false information that persists and influences later runs — is a recognised agentic risk (OWASP ASI06). Content that arrives from outside the trust boundary must never be written to durable memory without validation and provenance. Record, for every memory item: what it says, where it came from, when, and under whose authority.

Memory is a privacy obligation. Personal data in an agent’s memory is personal data. Subject access, rectification and erasure rights apply to it. If you cannot enumerate what your agent remembers about a named individual and delete it, you have a compliance defect regardless of how well the agent performs.

Business example

A client-service agent remembered every stated client preference indefinitely, including a two-year-old instruction to “always copy in Sarah”. Sarah had left. The agent kept adding her, drafts kept bouncing, and trust in the system fell. The fix was not technical sophistication but governance: preference memories carry a source, a date and a review interval; the account manager sees the remembered preferences on the client record and can correct them in one click. Memory without an owner becomes debt.

Common mistakes

  • Building RAG when a database query would answer the question exactly.
  • Storing everything because storage is cheap. The cost is not storage; it is retrieval noise and privacy exposure.
  • No provenance on memories, so nobody can tell what is true.
  • Never testing retrieval in isolation. Evaluate retrieval quality separately from answer quality — otherwise you cannot tell which half is broken.
  • Treating vector search as magic rather than as an approximation with a measurable hit rate.

Expert insight

Retrieval quality caps system quality. If the right chunk is not retrieved, no amount of model capability rescues the answer; if it is retrieved, a mid-tier model usually suffices. Given a fixed improvement budget on a RAG system, spend it on retrieval before you spend it on the model.

Knowledge check

  1. Distinguish knowledge from memory, and give one governance implication of the difference.
  2. Name the five memory types and give a business example of each.
  3. List three distinct RAG failure modes and their fixes.
  4. Why do production systems combine keyword and semantic search?
  5. What is memory poisoning and what control prevents it?
  6. What must every persisted memory item carry alongside its content?
  7. When is a knowledge graph the better choice than a vector store?

Challenge

Take twenty real questions your team answers from documents. Build or configure a retrieval system and measure only one thing: for each question, is the correct source in the top five results? Report the hit rate before writing a single line of answer-generation logic. Teams routinely discover their retrieval hit rate is nearer sixty per cent than the ninety-five per cent they assumed.