Design a grounded answer pipeline
Separate retrieval, answer construction, and evidence checking.
- Describe the stages of retrieval-augmented generation
- Preserve passage identifiers
- Return no answer when evidence is missing
Separate the responsibilities
A retrieval-augmented generation pipeline retrieves relevant context and supplies it to a generator. Indexing, retrieval, generation, and verification are distinct stages. If retrieval misses the right passage, a fluent generator cannot reliably reconstruct it.
Keep passage identifiers and source metadata through the pipeline. An answer should cite evidence that actually supports its claims, not merely include a plausible-looking link. A citation’s existence and its relevance are separate checks.
Start with extraction
Before connecting a generative model, build a system that returns a matching passage verbatim from your own knowledge base. This baseline costs no inference API calls and makes retrieval errors obvious. It is retrieval, not a full generative RAG system.
Define a no-answer path and evaluate it with questions outside the knowledge base. If you later add generation, bound the context size, treat retrieved text as untrusted data, and verify both claim support and cost.
A small experiment you can run.
This intentionally simple retriever returns stored text and its ID. It does not generate a new answer or claim that keyword overlap proves relevance.
documents = {"setup": "Create a virtual environment before installing packages.",
"splits": "Keep the final test set separate from model selection."}
def retrieve(question):
query = set(question.lower().strip("?.!").split())
ranked = sorted(documents, key=lambda key: -len(query & set(documents[key].lower().split())))
best = ranked[0]
overlap = len(query & set(documents[best].lower().split()))
return {"source_id": best, "passage": documents[best]} if overlap else {"status": "no_match"}
print(retrieve("Why keep a test set separate?"))
print(retrieve("Galactic weather tomorrow?"))
Save the file, open your terminal in that folder, and run python design-a-grounded-answer-pipeline.py. Use python3 or py if required by your installation. Setup guide
The first query returns the splits passage; the second returns no_match.
Design a grounded response contract.
- Require source IDs for supported answers.
- Allow an explicit no-answer status.
- Write an evaluation question that shares keywords but asks for an unsupported fact.
Compare with a suggested solution
Use a response with status, answer, and source_ids. Validate that IDs exist, then separately assess whether the passages entail the answer. A question can mention test sets while asking about a specific experiment absent from the documents.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python functools module