How can each token decide which other tokens matter?
Start here
How can each token decide which other tokens matter?
Attention computes content-dependent weights between elements, allowing each token to combine information from relevant context rather than relying only on fixed local recurrence.
Building interactive view…
Understand
Build the mental model
Attention computes content-dependent weights between elements, allowing each token to combine information from relevant context rather than relying only on fixed local recurrence. Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1
Queries / keys / values
For the “Queries / keys / values” stage, identify the incoming object, the rule applied to it, the state change produced, and the evidence that would reveal a mistake. Technical context for Attention & Transformers: Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order.
Practitioner checkpoint: Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.
What happens if…?
Break the assumption deliberately
Increase attention sharpness and watch one context token dominate the representation.
Move the control and explain what you expect before reading the visual.
Technical questionUse a tiny case to make the mechanism observable. Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Increase attention sharpness and watch one context token dominate the representation.
Practitioner lens
Use it responsibly
Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.
Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations. Then explain what should change if you deliberately test: Increase attention sharpness and watch one context token dominate the representation.
Worked exploration
Use the visual as an experiment, not decoration
For three tokens, compute query–key similarity scores for one token, apply softmax and form a weighted sum of value vectors. Change one similarity score and watch attention mass shift toward that token.
Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.
Prediction before interaction
Increase attention sharpness and watch one context token dominate the representation.
Exploration walkthrough
Turn the interaction into an evidence trail
For three tokens, compute query–key similarity scores for one token, apply softmax and form a weighted sum of value vectors. Change one similarity score and watch attention mass shift toward that token. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.
Record one observable quantity before the interaction and the same quantity afterwards.
Change one factor at a time so the causal effect of the control is inspectable.
Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Static orientation diagram for Attention & Transformers; use the interactive visual above to test how the relationships change.
Reference depth
Open the complete material
The flagship experience is the map. These pages contain the roads.