A knowledge assistant can return a confident answer and still be wrong because the document was read badly before retrieval began. A table may lose its row relationships. A heading may be detached from the section below it. A scanned page may produce incomplete text. If the source representation is damaged, a better prompt cannot repair it.
Retrieval-augmented generation, or RAG, means a model searches a collection of documents and uses the retrieved material when it answers. The useful question is not whether a model can answer a sample question. It is whether the system can find the right version, preserve the relevant structure, respect access rules, and show where the answer came from.
Start with a representative document set
Do not begin with a folder of tidy text files. Build a small test set that resembles the information the assistant will actually use. Include:
- A digitally generated PDF with headings, footnotes, and a table.
- A scanned document that requires OCR.
- A Word document with meaningful headings, lists, and a revision date.
- A presentation where the slide title changes the meaning of the bullets.
- Two revisions of the same procedure, with one clearly superseded.
- A document that a test user may see and another that the user must not see.
Use synthetic or licensed samples. A test set does not need real customer records to expose extraction failures. It does need enough variety to reveal what happens when layout, version, or permission changes.
Inspect the extracted representation
Before choosing a model, inspect what the ingestion step produces. Read the extracted text. Then compare it with the source document.
For each sample, check:
- Are headings retained and attached to the correct section?
- Does reading order follow the page or slide as a person would read it?
- Do table headers stay connected to the right values?
- Can the system identify page, slide, section, or paragraph anchors?
- Are dates, revision labels, and document identifiers preserved?
- Does OCR make important terms look uncertain or incomplete?
Vendor documentation can describe the available output without proving that your own files will be read correctly. OpenAI’s file search documentation describes hosted parsing, chunking, embeddings, and vector and keyword search. Google Document AI documents OCR output with text and layout information, as well as table and form extraction. Those capabilities are useful starting points, but your acceptance test still needs to inspect the result from your own representative files.
Test retrieval with questions that can fail
A retrieval test should contain more than questions with obvious answers. Write a set of answerable questions that depend on different document features:
- A question answered by a paragraph under a specific heading.
- A question that requires matching a table row to its column header.
- A question whose answer appears only on one slide.
- A question that requires the current revision rather than the older one.
- A question whose answer is absent from the collection.
- A question that a user is not allowed to answer from a restricted file.
For each question, record the expected source, the expected anchor, and the acceptable answer boundary. If the answer is not in the collection, the correct result is a clear statement that the source set does not answer it. A system that fills the gap with a plausible sentence has failed the test even if the sentence sounds useful.
Make revisions and permissions part of the test
Document assistants often fail at the edges of the collection. A superseded procedure may rank above the current one. A restricted document may leak into a result because the search index does not apply the user’s permissions. A draft may be treated as final because its filename looks authoritative.
Give every test document metadata that supports a real decision: owner, status, effective date, revision, audience, and source location. Then test the same question as different users. Confirm that the assistant cannot retrieve content the user is not allowed to read. Also test what happens when a document is withdrawn or replaced. Removing a file from a source folder does not prove that every indexed copy has disappeared.
Keep version selection visible. The answer should identify the document and revision it used, or link to the relevant page or section. A citation that points only to a large PDF makes review harder when the source contains several similar instructions.
Use a small acceptance scorecard
Turn the test set into a repeatable scorecard before connecting business data. Record the result for every question, not only the ones that went well.
- Was the expected document retrieved?
- Was the expected page, slide, section, or table region identified?
- Did the answer preserve the source’s limits?
- Did the system refuse questions with no supported answer?
- Did permission filtering remove restricted material?
- Did the current revision outrank a superseded one?
- Could a reviewer reproduce the answer from the displayed source anchor?
Separate extraction errors from retrieval errors and answer errors. If a table was damaged during extraction, changing the retrieval model will not solve the underlying problem. If the correct passage was retrieved but the answer ignored a qualification, the next investigation belongs in the answer-generation step.
Assign ownership before the pilot
Someone must own the documents, the ingestion process, and the acceptance test. The owner should know who can approve a replacement, how a superseded file is marked, and what happens when a source changes. A technical team can operate the pipeline, but it should not silently decide which business procedure is authoritative.
Start with a narrow collection and a documented test set. Keep the first pilot read-only if the assistant will influence operational decisions. Require a human to review answers that affect customers, money, safety, access, or contractual commitments. The point is not to prove that a model never makes a mistake. The point is to find the failure modes while the collection and workflow are still small enough to correct.
Before connecting an assistant to your documents, ask greenstar technology to review the information structure and retrieval requirements. A focused review can expose weak source material before it becomes a trusted-looking answer.
Source lead: archived X post about document ingestion problems. The post supplied the editorial starting point, not independent validation of a named tool or feature.