Learn AI/Language AI
LESSON 22 / 36Intermediate 40 min with practice

Build a TF–IDF document search

Retrieve relevant passages using an inspectable lexical baseline.

WHAT YOU WILL LEARN
  • Weight common and rare terms
  • Compare query and document vectors
  • Recognize lexical retrieval limitations

Not every word is equally useful

Term frequency measures how often a term appears in a document. Inverse document frequency gives more weight to terms appearing in fewer documents. Multiplying the two produces a lexical representation that can highlight distinctive words.

Different implementations use different smoothing and normalization conventions. The exercise uses raw counts and log((1+N)/(1+document_frequency))+1, then compares vectors with cosine similarity. Keep the convention fixed when comparing results.

Retrieve evidence, not an answer

Retrieval finds candidates. It does not establish that a passage contains a correct answer or that an answer generated from it is grounded. Exact-word retrieval can miss synonyms and find misleading keyword matches.

Start with a lexical baseline because it is cheap and inspectable. Evaluate relevant passage retrieval on questions written separately from the indexed documents. Later you can compare dense representations or combine both approaches.

PUT THE IDEA INTO CODE

A small experiment you can run.

All corpus documents are indexed for retrieval. This is different from fitting a supervised model on evaluation labels; keep evaluation questions and relevance judgments separate.

tf-idf-document-search.py
import math
from collections import Counter
docs = ["python functions organize code", "neural networks learn patterns", "python code trains models"]
tokens = [doc.split() for doc in docs]
vocab = sorted(set(sum(tokens, [])))
idf = {w: math.log((1+len(docs))/(1+sum(w in d for d in tokens)))+1 for w in vocab}
def vector(text):
    counts = Counter(text.lower().split())
    return [counts[w]*idf[w] for w in vocab]
def cosine(a, b):
    denom = math.sqrt(sum(x*x for x in a)*sum(x*x for x in b))
    return sum(x*y for x, y in zip(a, b))/denom if denom else 0.
query = vector("python code")
ranked = sorted(enumerate(docs), key=lambda row: -cosine(query, vector(row[1])))
for index, doc in ranked:
    print(index, round(cosine(query, vector(doc)), 3), doc)
Copy code

Save the file, open your terminal in that folder, and run python tf-idf-document-search.py. Use python3 or py if required by your installation. Setup guide

What to expect

Python-related documents rank ahead of the neural-network-only document.

YOUR TURN

Test a synonym failure.

  1. Search for programming instead of python code.
  2. Observe the query vector when all terms are unknown.
  3. Add an explicit no-match result rather than returning the first document as a confident answer.
Compare with a suggested solution

The query becomes a zero vector and all scores are zero. Require a positive score for lexical overlap, and use a validation set to choose any stronger confidence threshold. Adding synonyms can help but must be evaluated.

CHECK YOUR UNDERSTANDING

One idea to take with you.

What does TF–IDF primarily match?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

scikit-learn: text feature extractionAttention Is All You Need: original paper