Learn AI/Language AI
LESSON 24 / 36Intermediate 40 min with practice

Attention in small numbers

Compute attention weights and see what is being mixed.

WHAT YOU WILL LEARN
  • Compare a query with keys
  • Normalize attention scores
  • Mix value vectors

Queries ask, keys match, values contribute

In attention, a query is compared with keys to produce scores. Softmax turns those scores into weights, and a weighted sum mixes the corresponding values. Scaled dot-product attention divides scores by the square root of the key dimension.

In a transformer, learned projections produce queries, keys, and values from representations. Our example provides the vectors directly so each multiplication can be inspected. Attention can be used inside many architectures; one weighted sum is not a full language model.

Attend within the allowed context

A causal language model masks future positions so a token cannot access the continuation it is supposed to predict. Padding masks prevent artificial padding positions from receiving attention.

Attention weights describe one internal computation. They are not a complete explanation of a model’s reasoning, nor do they prove a source is factually reliable. Other layers and transformations also contribute to the output.

PUT THE IDEA INTO CODE

A small experiment you can run.

The most aligned key receives the largest weight. The output mixes values, not keys, which can represent different information.

attention-in-small-numbers.py
import math
query = [1., 0.]
keys = [[1., 0.], [0., 1.], [-1., 0.]]
values = [[2., 0.], [0., 2.], [1., 1.]]
scores = [sum(q*k for q, k in zip(query, key))/math.sqrt(len(query)) for key in keys]
exps = [math.exp(score-max(scores)) for score in scores]
weights = [value/sum(exps) for value in exps]
output = [sum(w*v[d] for w, v in zip(weights, values)) for d in range(2)]
print("Attention weights:", [round(w, 3) for w in weights])
print("Mixed value:", [round(v, 3) for v in output])
Copy code

Save the file, open your terminal in that folder, and run python attention-in-small-numbers.py. Use python3 or py if required by your installation. Setup guide

What to expect

All weights are positive, sum to one, and the first key gets the largest weight.

YOUR TURN

Apply a causal-style mask.

  1. Exclude the third key/value position before softmax.
  2. Recalculate the two remaining weights.
  3. Explain why masking must happen before normalization.
Compare with a suggested solution

Remove masked scores or assign negative infinity before softmax. Setting a weight to zero afterward without renormalizing leaves the weights summing to less than one and changes the intended operation.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Which vectors are combined to form the attention output?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

scikit-learn: text feature extractionAttention Is All You Need: original paper