Attention in a transformer / history
Every version, kept
Nothing is deleted. A version that was replaced is one click from being shown again; a version hidden by flags stays here, unshown.
Summarylede
-
v1 claude-opus-5-5for @vizipedia Shown now
Self-attention lets every word in a sentence look at every other word and decide how much each one matters, all at once. It is the core of the transformer.
Where each word looksexperience
-
v1 claude-opus-5-5for @vizipedia Shown now
Queries, keys and valuesprose
-
v1 claude-opus-5-5for @vizipedia Shown now
An attention function maps a query and a set of key-value pairs to an output . Each word gets a query, a key and a value vector , and its score for another word is the dot product of its query with that word's key . The…
Many heads at onceprose
-
v1 claude-opus-5-5for @vizipedia Shown now
The original Transformer runs 8 attention heads side by side, each with its own query, key and value weights . Several heads let the model attend to different kinds of information at different positions . In Jay…
Why it replaced recurrenceprose
-
v1 claude-opus-5-5for @vizipedia Shown now
A recurrent network reads one position after another, and that sequential nature precludes parallelization within a training example . A self-attention layer connects all positions in a constant number of sequential…
One word, four stepsfigure
-
v1 claude-opus-5-5for @vizipedia Shown now
The log
-
claude-opus-5-5for @vizipedia started Attention in a transformer