Module 6 of 12
Knowledge, Documents and Data
How do you ground AI in your own information without drowning it?
Learning objectives. Provide an AI with documents, data, and knowledge sources correctly; decide what to include and exclude from a knowledge source; handle conflicting information without letting the AI silently resolve it; and understand grounding as the primary defence against hallucination.
Why this matters
Most professional AI use eventually involves feeding it your own material — reports, policies, spreadsheets, research, notes. Done well, this is where AI becomes genuinely reliable for real work. Done carelessly, it's where hallucination does the most damage, because a confident, fluent, wrong answer that looks grounded in your own documents is more dangerous than an obviously made-up one.
Plain-English explanation
Grounding means basing an AI's answer on real, supplied material rather than its generalised training. When you give an AI your actual pricing sheet and ask it to answer a pricing question, you're grounding the answer. When you don't, and it answers from general knowledge about "typical" pricing, you get a plausible-sounding guess dressed as an answer.
Core lesson
PLAIN ENGLISH — what to include, what to exclude
Include: the specific document, policy, dataset, or notes that actually contain the answer. Current versions only — Module 5's context principle applies here directly: one relevant, current policy beats five superseded versions "for completeness."
Exclude: anything outdated, contradictory-and-unresolved, or irrelevant to the specific question. A knowledge base is not an archive; it's a curated set of what should currently be trusted.
How to handle conflicting information: never let the AI silently pick a winner between two contradictory sources. Explicitly instruct it: "if the sources disagree, state the disagreement rather than resolving it" — Module 5's synthesis technique note applies here directly, and it's one of the most common places hallucination-adjacent errors happen unnoticed.
How to prioritise authoritative sources: tell the AI, explicitly, which source wins if there's a conflict — "the current staff handbook overrides anything in old email threads" — rather than assuming it will infer your intended hierarchy.
PRACTITIONER — the practical grounding workflow
- Identify the specific document or data that actually answers the kind of question you're asking — not "everything we have," the relevant material.
- Supply it directly (paste, upload, or — in a Level 3 RAG setting, covered next — retrieve it automatically).
- Instruct the AI explicitly to answer from the supplied material, and to say when something isn't covered by it, rather than filling the gap from general knowledge.
- Ask for citations or direct quotes where the claim came from, when the stakes justify it — this makes verification (Module 11) fast instead of a re-read of the whole source.
- Spot-check: for anything consequential, actually confirm the AI's claimed source location says what it claims.
ADVANCED — retrieval-augmented generation (RAG), briefly
At scale — too much material to paste every time — the standard architecture is retrieval-augmented generation (RAG): documents are broken into chunks, indexed so relevant chunks can be found by meaning (not just exact keyword match), and the most relevant chunks are retrieved and inserted into context automatically for each query, ideally with citations. The foundational paper is Lewis et al. (2020). Current practice has moved past naive RAG toward hybrid search — combining semantic (meaning-based) and keyword search, since semantic search alone often misses exact identifiers like invoice numbers or product codes — and toward evaluating retrieval quality separately from answer quality, because if the right material was never retrieved, no amount of downstream cleverness rescues the answer. This is developed further in Module 7; Level 3 learners building this themselves should read the companion Holistic Agent course's Agent Memory and Knowledge module in full.
A structured-data rule worth stating plainly: if the answer to a question lives in a spreadsheet or database as an exact row, query that source directly rather than routing it through a document-search system built for unstructured text. RAG over a PDF export of a table you could have queried directly is a common, entirely avoidable source of approximate answers to questions that have an exact one.
Beginner example
Ungrounded: "What's a fair returns policy for an online clothing shop?" — produces a plausible generic answer that may not match your actual policy.
Grounded: pasting your actual returns policy and asking "using this policy, write a response to a customer asking to return a worn item after 45 days" — the answer is now anchored to your real rules, not a generic industry norm.
Business example
A small law firm's paralegal used AI to summarise case files, initially pasting entire multi-hundred-page bundles and getting summaries that missed key details buried in the middle (Module 5's finding, showing up in a real workflow). Rebuilt: summarise section by section, with the specific clause type flagged before each pass ("summarise the liability clauses in this section specifically"), and cross-check names and dates against the source rather than trusting the summary's restatement of them. Slower per document, meaningfully more accurate, and — critically — the firm's review process now catches the rare miss before it reaches a client, because verification was built into the workflow rather than assumed.
Practical exercise — add useful context, then remove the irrelevant
Take a real question you'd ask about your own work that depends on a specific document (a policy, a proposal template, a dataset). First, answer it with the AI with no supplied document — note the generic, ungrounded answer. Then supply the actual relevant excerpt (not the whole document) and re-ask. Compare specifically for accuracy, not just tone.
Common mistakes
- Assuming an AI "knows" your business because you've mentioned it in conversation before — it doesn't, unless the actual material is in context now.
- Letting the AI silently resolve contradictory sources instead of flagging the conflict.
- Feeding an entire document when a specific excerpt would ground the answer just as well, at a fraction of the context cost (Module 5).
- Treating extraction or summarisation from a real document as automatically accurate, without spot-checking against the source (Module 4's extraction failure mode).
Expert insight
The instinct to "give it everything" comes from a good place — wanting the AI to have what it needs. But per Module 5's evidence, the instinct that actually produces better, more grounded answers is the opposite: give it exactly what answers this question, tell it explicitly to stick to that material, and tell it what to do when the material doesn't cover something. Grounding is not about volume. It's about precision.
Knowledge check
- Define grounding in one sentence.
- What should you explicitly instruct the AI to do when two supplied sources disagree?
- Why is querying a database directly usually better than running it through document-style retrieval?
- Name two things a knowledge source should exclude, not just what it should include.
- What's the fastest way to make verification of a grounded answer practical rather than a full re-read?
Module summary
Grounding — basing answers on real, supplied material rather than generalised training — is the primary defence against hallucination for anything that matters. What you include, what you exclude, how you handle conflicting sources, and how you prioritise authority all determine whether "grounded" actually means accurate, or just means confident with extra steps.
Further exploration
Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv:2005.11401 (2020) — the origin paper for RAG, developed further in Module 7.