Deep Learning · Flagship experience

Attention & Transformers

How can each token decide which other tokens matter?

Start here

How can each token decide which other tokens matter?

Attention computes content-dependent weights between elements, allowing each token to combine information from relevant context rather than relying only on fixed local recurrence.

Building interactive view…
Understand

Build the mental model

Attention computes content-dependent weights between elements, allowing each token to combine information from relevant context rather than relying only on fixed local recurrence. Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.

Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Queries / keys / values

For the “Queries / keys / values” stage, identify the incoming object, the rule applied to it, the state change produced, and the evidence that would reveal a mistake. Technical context for Attention & Transformers: Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order.

Practitioner checkpoint: Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.
What happens if…?

Break the assumption deliberately

Increase attention sharpness and watch one context token dominate the representation.

Move the control and explain what you expect before reading the visual.

Technical lens

Formalise what the visual is doing

Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order.

Technical questionUse a tiny case to make the mechanism observable. Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Increase attention sharpness and watch one context token dominate the representation.
Practitioner lens

Use it responsibly

Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.

Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations. Then explain what should change if you deliberately test: Increase attention sharpness and watch one context token dominate the representation.
Worked exploration

Use the visual as an experiment, not decoration

For three tokens, compute query–key similarity scores for one token, apply softmax and form a weighted sum of value vectors. Change one similarity score and watch attention mass shift toward that token.

Technical lens

Scaled dot-product attention uses softmax(QKᵀ/√d)V. Multi-head attention learns multiple relation subspaces; positional information restores order.

Practitioner check

Inspect sequence length/memory costs and remember attention weights are not automatically causal explanations.

Prediction before interaction
Increase attention sharpness and watch one context token dominate the representation.
Exploration walkthrough

Turn the interaction into an evidence trail

For three tokens, compute query–key similarity scores for one token, apply softmax and form a weighted sum of value vectors. Change one similarity score and watch attention mass shift toward that token. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.

  • Record one observable quantity before the interaction and the same quantity afterwards.
  • Change one factor at a time so the causal effect of the control is inspectable.
  • Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Visual demonstration of Attention & Transformers
Static orientation diagram for Attention & Transformers; use the interactive visual above to test how the relationships change.
Reference depth

Open the complete material

The flagship experience is the map. These pages contain the roads.

Continue this exact concept

Choose depth, practice or application.

These destinations are explicitly mapped to Attention & Transformers; they are not generic landing-page fallbacks.