Back to News & insightsEngineering

From a question to evidence: watch a grounded answer take shape

Trace retrieval, inspect sources, and remove the evidence.

Editorial guide · Updated October 3, 2026 · 8 min read
Three fictional policy documents pass through an evidence check to a supported twelve-day answer.

A retrieval-augmented assistant does not become trustworthy merely because a search step sits before a language model. The search can return an obsolete rule, the context can omit an exception, and the answer can claim more than the retrieved passage supports. To understand the system, follow one question and inspect what crosses each boundary.

This visual guide uses a deliberately fictional equipment-rental policy. The question is simple: how long may a member keep a standard loan? Three short documents provide a current rule, an older rule, and a rule for a different plan. Their values are invented for this teaching example. They are not a real organization's terms and are not evidence about a model's benchmark performance.

Scroll through the architecture diagram or choose a stage to follow the question from retrieval to verification. In the evidence exercise, remove the current document and inspect the response. The important change is not a more elaborate sentence. It is the system recognizing that the remaining evidence does not establish the current answer.

Follow the question, not a cloud of boxes

01 / FOLLOW THE ARCHITECTURE

One question, an inspectable evidence trail

Reduced motion is on. Choose a stage or move the slider to explore.

STAGE 01

Find candidate passages

Question: how long is a current standard equipment loan? Search can return all three fictional policies, including unsuitable ones.

Read every stage without animation
  1. Find candidate passages

    Question: how long is a current standard equipment loan? Search can return all three fictional policies, including unsuitable ones.

  2. Check which policy applies

    Exclude the archived seven-day rule and the twenty-one-day extended plan. Similar wording does not make either authoritative here.

  3. Preserve the evidence boundary

    Pass the current standard policy with its identifier and scope. Keep source text separate from application instructions.

  4. Express the supported conclusion

    The twelve-day answer comes from the fictional current policy. Do not add an unsupported renewal promise or late fee.

  5. Check the claim against its source

    The duration and plan must match the passage. Missing eligible evidence should produce an explicit knowledge boundary.

Original conceptual diagram. Animation speed is illustrative; no model is running and no measured latency is implied.

The request begins with a precise subject: a member asking about a standard loan under the current policy. That scope matters. A document describing an extended plan may contain similar words while answering a different question. A previous policy may be historically accurate while being wrong for a current request.

The animated path separates retrieval, eligibility checks, context assembly, drafting, and claim verification. Each stage receives an explicit input and produces an inspectable output. The illustration is a teaching architecture, not a recording of a deployed model thinking. It helps reveal where an application must make decisions that a generic search score cannot settle.

Retrieve candidates before calling them evidence

A search result is a candidate passage. It becomes useful evidence only after its meaning and applicability are checked against the question. Similarity can help locate material, but semantic proximity alone does not establish that a document is current, authorized for the reader, or relevant to the exact plan being discussed.

The original retrieval-augmented generation research combines a retrieval component with a generative model for knowledge-intensive tasks. That architecture motivates bringing external material into generation. It does not establish that every retrieved result is correct or that a citation automatically proves every sentence in the answer.

Read the original research paper on arXiv

In the fictional example, all three passages plausibly match a query about borrowing equipment. A naive system could choose the shorter old rule or the more generous plan simply because the wording looks relevant. Keeping the candidates visible lets the reader see why retrieval and selection must be treated as separate responsibilities.

Eligibility belongs outside the model's imagination

Suppose the archive records which policy is active and which plan each passage covers. Those fields should participate in the selection process. The application can exclude an archived policy from a current-policy answer and exclude an extended-plan passage from a standard-plan question. These are explicit rules over maintained records, not facts the generator should invent.

Access permission is another eligibility condition. A passage that the user may not read should not be sent into the user's answer context and then hidden only at the display stage. The appropriate access boundary depends on the application, but it should be enforced by trusted software using the actual user identity and resource policy.

The teaching diagram uses a simplified current-plan check rather than a full authorization system. That simplification is visible because each document has a named status. In a real implementation, preserve the identifiers and decision reasons needed to audit why a passage was included or excluded without exposing sensitive document text in unnecessary logs.

Assemble context that can still be traced

After selection, the generator needs enough context to answer the question and identify the supporting passage. A useful context package preserves the source identifier, the relevant excerpt, and the scope that makes it applicable. Stripping these details to save a few tokens can make later verification harder.

Do not concatenate every retrieved document without a clear boundary. A model may blend an old duration with a current exception or borrow a benefit from a different plan. Distinct source labels and concise excerpts help the application inspect the resulting claim. They do not guarantee perfect behavior, but they make failures less mysterious.

The example's current standard policy permits a twelve-day loan. The older policy says seven days, and the extended plan says twenty-one. These deliberately conflicting values force the system to use scope rather than majority voting or numerical generosity. Repeating the wrong document several times would not make it the right source.

Try the evidence removal experiment

02 / TRY THE FAILURE CASE

What happens when the evidence disappears?

How long may a standard member borrow equipment under the current policy?

SUPPORTED CLAIM

A standard member may borrow equipment for 12 days.

Support: Current standard policy
Fictional rental policies. This deterministic browser demonstration makes no AI request and is not a model evaluation.

The interactive evidence desk below exposes those three fictional documents. Toggle the current standard policy off. The deterministic teaching response should now state that the current duration cannot be established from the eligible evidence. It should not promote an archived passage or a different plan into authority merely to avoid leaving the user without a number.

Turn the current document back on and inspect the claim-to-source connection. The answer is supported by a particular passage, not by a vague declaration that the system searched the web. The exercise is intentionally small so that every choice can be checked directly by a human reader.

This widget runs entirely in the browser and makes no AI API call. Its output is a scripted demonstration of an evidence policy. That distinction matters: an example showing how an application should behave is not an evaluation proving that a language model will always follow the same rule.

Verify claims at the right level

An answer can contain several claims, each needing different support. In the fictional policy case, the duration is one claim and the plan's applicability is another. If the assistant adds a late fee, a renewal promise, or an exception absent from the source, the presence of a valid duration citation does not support those additions.

A useful verification step compares individual claims with the relevant passages and checks whether the evidence actually entails them. It should distinguish direct support, reasonable but uncertain inference, and unsupported content. The application can then revise the answer, request clarification, or hand the case to a person according to the task's requirements.

Automated verification is itself a component to evaluate. A second model can repeat the first model's mistake or approve a fluent unsupported sentence. Preserve reviewed examples where the verifier should reject a claim, and test the complete pipeline. Adding another box to the architecture diagram does not automatically add dependable assurance.

Keep instructions separate from retrieved material

Documents are data supplied to the application, not a new source of system authority. A retrieved page could contain text telling an assistant to ignore earlier instructions, reveal secrets, or contact an unrelated service. The application should not treat that text as permission to change its rules or perform an action.

The safe boundary is broader than wording a prompt politely. Restrict the actions and data available to the generation step, enforce access in trusted code, and validate any proposed operation before execution. Retrieval should not become a route by which arbitrary source text acquires control over the application.

OWASP's prompt-injection guidance describes this class of risk and emphasizes defenses across the system. The practical implication for this example is that a policy passage can support a statement about a loan; it cannot authorize a new tool call or override the rules for deciding which policy applies.

OWASP: Prompt injection risks and mitigations

Evaluate retrieval and answering separately

Create a reviewed set of questions with known supporting passages. First measure whether the retrieval stage brings those passages into the candidate set. Then inspect whether eligibility and context assembly retain the right material. Finally evaluate whether the answer expresses the supported conclusion without adding unsupported details.

This separation makes diagnosis much easier. If the relevant policy never reaches the candidate set, changing the final answer prompt may not address the underlying issue. If the correct passage is present but the generator chooses the archived rule, retrieval recall alone cannot explain the failure. Each boundary needs an observable outcome.

Include questions with no eligible answer. A pipeline that always produces a plausible sentence may appear helpful on easy examples while failing the most important absence test. Missing evidence should be represented in the evaluation contract, not treated as an inconvenient case to remove from the dataset.

Make changes propagate through the whole path

A source archive changes over time. A policy can be replaced, a document can be deleted, or a user's access can be revoked. The retrieval index, cached context, and generated answers need a plan for reflecting those changes. Updating the source file alone does not prove that every downstream copy has stopped serving the old material.

Track source versions and the dependencies of cached outputs where the application needs that precision. Test a policy replacement from ingestion through retrieval and answer delivery. Check whether a previously saved answer remains appropriately labeled as historical or is recomputed when the product promises a current answer.

The evidence-removal control is a miniature version of this operational test. It asks whether the system's confidence changes when its support disappears. A robust application should make that dependence explicit rather than preserve an old confident answer simply because it is convenient to cache.

Design the answer for inspection

Present the conclusion with a concise source link and enough context for the reader to verify its applicability. Avoid overwhelming the user with an unfiltered document dump. The goal is a usable explanation whose important claims can be checked, not a performance of research made from many unrelated citations.

For the fictional loan question, a strong answer gives the current standard duration and points to the current standard policy. If that evidence is unavailable, it says what could not be established and what additional record would resolve the question. The response remains useful because it describes the boundary of its knowledge precisely.

Follow the animated path once more with that output in mind. Retrieval finds candidates, eligibility narrows their authority, context preserves traceability, generation expresses a conclusion, and verification checks its support. The value comes from the connections between those stages. That is what turns a collection of documents and a language model into an answer a reader can inspect.

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.