AI Systems Research guide

TRACE vs RAG: one fetches the text, the other checks what it proves

An answer can quote the right document and still draw the wrong conclusion.

Split basement archive: on the left a finger pointing at a filing-cabinet drawer, on the right a single blank page under a clamp lamp
Finding the page is one job. Reading what it proves is another.

Direct answer

RAG (retrieval-augmented generation) is a software design: before an AI model answers, it fetches passages from a document collection and writes from them. TRACE is a research method from The Verified Research Loop: Target, Retrieve, Assess, Cross-check, Export. RAG answers the question 'which text did the model use?' TRACE answers 'what does the evidence prove, and may we act on it?' RAG can feed TRACE's Retrieve step. Neither one proves a conclusion true on its own.

RAG and TRACE solve different problems. RAG is a software design that fetches text for an AI model before it answers. TRACE is a research method that checks what the sources prove before someone acts. RAG works inside a product, on every question. TRACE works around one decision. You can use them together, and most of the confusion comes from assuming one does the other’s job.

What is RAG, in plain words?

RAG stands for retrieval-augmented generation. When you ask a question, the system first searches a document collection (a company wiki, a policy library, a folder of contracts), picks the passages that look most relevant, and hands them to a language model, which writes the answer from them.

The name comes from a 2020 research paper by Patrick Lewis and colleagues, which paired a text generator with a searchable index of Wikipedia. What is RAG? covers the method in more depth.

What is TRACE, in plain words?

TRACE is a five-step method for checking research before a decision. It comes from The Verified Research Loop by Len P. van der Hof, available in English and Dutch. You do not need the book to use it.

  • Target: tie the question to a decision and state the exact claim you need to test.
  • Retrieve: obtain the actual source documents, not snippets or summaries.
  • Assess: judge what each source can prove for this claim. Sources get one of four roles (Primary, Secondary, Advocacy, or Grey), and none of the roles is a verdict.
  • Cross-check: compare the sources claim by claim and trace each one to its origin. Three articles repeating one press release count as one origin.
  • Export: hand over a dated memo that separates what is supported from what is still open, with an owner for every gap.

What is the difference, side by side?

RAGTRACE
What it isA software designA research method
Who runs itA system, automatically, on every questionA person (with or without AI help), for one decision
Main questionWhich passages should the model see?What do the sources prove, and may we act on it?
OutputA generated answer, ideally with citationsA claim ledger and a dated decision memo
Typical failureMissing, stale, or wrong passage; answer drifts from its passageClaim broader than the evidence; one origin counted as several
Done whenEach sentence is supported by the passages the model receivedEvery important claim has a status, and every gap has an owner
Framework on this siteGRAINTRACE

How can the document be right and the conclusion wrong?

Here is a hypothetical. An operations team asks its internal assistant: “What does the incident report say caused Tuesday’s outage?” The RAG system retrieves the current report and answers correctly, quoting the right passage: a timeout at an external payment provider. Retrieval did its job.

A week later, someone writes in a planning document: “This provider causes most of our failed payments.” The author cites the same answer and proposes switching providers.

The citation is real. The passage is accurate. But one incident report cannot tell you what causes most failures. That claim needs a defined period, a count of all failed payments, the cause of each, and a look at evidence that points elsewhere.

RAG was never asked that question, so it could not fail it. TRACE is where the broader claim gets tested:

  1. Target states the real question: “Share of failed payments by cause, last 90 days.”
  2. Retrieve collects the payment logs and all incident reports for the period, not only Tuesday’s.
  3. Assess notes that Tuesday’s report is primary evidence for one incident only.
  4. Cross-check finds no source for “most”.
  5. Export records the incident finding as supported, the “most failures” claim as Unresolved, and puts the switch on HOLD, the TRACE word for “this action waits until a named person closes the gap or formally accepts the risk”.

Which one do I need?

  • You are building an assistant that answers from your own documents: RAG, built and debugged with GRAIN.
  • You must decide something from outside sources (a market, a competitor, a regulation, a supplier): TRACE. AI tools, including RAG systems, can help with the Retrieve step.
  • People act on your RAG assistant’s answers: both. RAG gets the right passage into the answer. A TRACE-style check decides whether that answer supports the action.
  • You have one document and one question: neither. Open the document and read it. When to use RAG covers the cases where building RAG is the wrong job.

Where does GRAIN fit?

GRAIN is the matching framework on the RAG side. It comes from The RAG Engineer, also available in English and Dutch, and names the five places a retrieval system can break:

  • Gather corpora: which documents may enter the searchable collection, with their permissions and versions.
  • Rank and rerank: finding candidate passages, then putting them in order for the actual question.
  • Assemble context: choosing which passages go to the model, in what order, within its input limit.
  • Inspect failures: tracing a wrong answer back to the step that caused it.
  • Navigate freshness: keeping the collection up to date and showing how old an answer could be.

What is GRAIN? has more. In one line: GRAIN checks the pipe, TRACE checks the claim.

When an answer fails, whose problem is it?

Name the broken step before you fix anything.

What you seeWhere it brokeWhose job
The answer relies on an outdated reportThe collection was not updatedRAG side: Navigate freshness
The right report was retrieved, but the answer drops a conditionThe sentence does not match its passageRAG side: Assemble context and answer checks
Several citations all go back to one accountThe sources share one originTRACE: Cross-check
An unresolved claim is used to justify a decisionThe decision outran the evidenceTRACE: Export, with an owner and HOLD

What do people get wrong?

  • “Our AI cites its sources, so the answer is verified.” A citation is a route back to the evidence. Someone still has to check that the passage supports the exact sentence, including its scope and date. What is LLM grounding? explains that check.
  • “TRACE needs an AI system.” It does not. A browser and a spreadsheet are enough.
  • “RAG makes the answer true.” RAG makes the answer depend on specific passages. That is what lets you check it, not a guarantee that it is right.
  • “Finishing TRACE proves the conclusion.” It shows what the evidence supports and where it stops. The book is explicit that TRACE has not been shown to make research more accurate or faster.

Try this today (10 minutes)

Take one AI answer your team has acted on, or is about to. Write three lines under it:

  1. Passage: which text did the answer come from, and can you open it?
  2. Sentence: does that text support the sentence exactly, including its scope and date?
  3. Decision: does the sentence support the action being proposed?

The first “no” tells you where to work. Lines 1 and 2 belong to whoever runs the RAG system. Line 3 belongs to whoever owns the decision, and that is where TRACE starts.

Cite this:TRACE vs RAG: one fetches the text, the other checks what it proves.Len P. van der Hof. https://lenvanderhof.com/en/blog/trace-vs-rag/ ·

Terminology

Sources

  1. What is RAG? When AI looks up a source before it answers
  2. What is GRAIN? The inspectable retrieval loop
  3. What is LLM grounding? The sentence has to open a folder
  4. The Verified Research Loop
  5. The RAG Engineer
  6. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks · arXiv (Lewis et al., 2020)

Further reading

Markdown for LLMs