Build a TF–IDF document search
Retrieve relevant passages using an inspectable lexical baseline.
- Weight common and rare terms
- Compare query and document vectors
- Recognize lexical retrieval limitations
Not every word is equally useful
Term frequency measures how often a term appears in a document. Inverse document frequency gives more weight to terms appearing in fewer documents. Multiplying the two produces a lexical representation that can highlight distinctive words.
Different implementations use different smoothing and normalization conventions. The exercise uses raw counts and log((1+N)/(1+document_frequency))+1, then compares vectors with cosine similarity. Keep the convention fixed when comparing results.
Retrieve evidence, not an answer
Retrieval finds candidates. It does not establish that a passage contains a correct answer or that an answer generated from it is grounded. Exact-word retrieval can miss synonyms and find misleading keyword matches.
Start with a lexical baseline because it is cheap and inspectable. Evaluate relevant passage retrieval on questions written separately from the indexed documents. Later you can compare dense representations or combine both approaches.
A small experiment you can run.
All corpus documents are indexed for retrieval. This is different from fitting a supervised model on evaluation labels; keep evaluation questions and relevance judgments separate.
import math
from collections import Counter
docs = ["python functions organize code", "neural networks learn patterns", "python code trains models"]
tokens = [doc.split() for doc in docs]
vocab = sorted(set(sum(tokens, [])))
idf = {w: math.log((1+len(docs))/(1+sum(w in d for d in tokens)))+1 for w in vocab}
def vector(text):
counts = Counter(text.lower().split())
return [counts[w]*idf[w] for w in vocab]
def cosine(a, b):
denom = math.sqrt(sum(x*x for x in a)*sum(x*x for x in b))
return sum(x*y for x, y in zip(a, b))/denom if denom else 0.
query = vector("python code")
ranked = sorted(enumerate(docs), key=lambda row: -cosine(query, vector(row[1])))
for index, doc in ranked:
print(index, round(cosine(query, vector(doc)), 3), doc)
Save the file, open your terminal in that folder, and run python tf-idf-document-search.py. Use python3 or py if required by your installation. Setup guide
Python-related documents rank ahead of the neural-network-only document.
Test a synonym failure.
- Search for programming instead of python code.
- Observe the query vector when all terms are unknown.
- Add an explicit no-match result rather than returning the first document as a confident answer.
Compare with a suggested solution
The query becomes a zero vector and all scores are zero. Require a positive score for lexical overlap, and use a validation set to choose any stronger confidence threshold. Adding synonyms can help but must be evaluated.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
scikit-learn: text feature extractionAttention Is All You Need: original paper