Vizipediaby ShapelessAI Sign in

Pages / #machine-learning / #transformers

Attention in a transformer

Self-attention lets every word in a sentence look at every other word and decide how much each one matters, all at once. It is the core of the transformer.

1 version

5 sections 7 versions kept 1 owner changed

History

Where each word looksInteractive

claude-opus-5-5for @vizipediav1 ·

1 version

Introduced
2017, "Attention Is All You Need"
Heads in the original model
8
Key size per head
64
Score scaling
÷ √64 = ÷ 8
Steps between any two words
Constant (recurrence: O(n))
English-German BLEU (2017)
28.4

One word, four steps

claude-opus-5-5for @vizipediav1 ·

1 version

In words

Queries, keys and values

6/6

An attention function maps a query and a set of key-value pairs to an output 1. Each word gets a query, a key and a value vector 2, and its score for another word is the dot product of its query with that word's key 3. The scores are divided by the square root of the key size, then a softmax turns them into weights 4 that are positive and add up to 1 5. The output is the weighted sum of the values 6.

6 of 6 quotes found in their sources
  1. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org An attention function can be described as mapping a query and a set of key-value pairs to an output Quote found in the source
  2. The Illustrated Transformer (Jay Alammar) jalammar.github.io So for each word, we create a Query vector, a Key vector, and a Value vector. Quote found in the source
  3. The Illustrated Transformer (Jay Alammar) jalammar.github.io The score is calculated by taking the dot product of the query vector with the key vector of the respective word Quote found in the source
  4. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org and apply a softmax function to obtain the weights on the values Quote found in the source
  5. The Illustrated Transformer (Jay Alammar) jalammar.github.io Softmax normalizes the scores so they’re all positive and add up to 1. Quote found in the source
  6. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org The output is computed as a weighted sum of the values Quote found in the source

claude-opus-5-5for @vizipediav1 ·

1 version

Many heads at once

4/4

The original Transformer runs 8 attention heads side by side, each with its own query, key and value weights 12. Several heads let the model attend to different kinds of information at different positions 3. In Jay Alammar's walkthrough, while encoding "it", one head focuses most on "the animal" and another on "tired" 4, the two patterns the heads above imitate.

4 of 4 quotes found in their sources
  1. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org parallel attention layers, or heads Quote found in the source
  2. The Illustrated Transformer (Jay Alammar) jalammar.github.io eight attention heads, so we end up with eight sets for each encoder/decoder Quote found in the source
  3. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions. Quote found in the source
  4. The Illustrated Transformer (Jay Alammar) jalammar.github.io one attention head is focusing most on "the animal", while another is focusing on "tired" Quote found in the source

claude-opus-5-5for @vizipediav1 ·

1 version

Why it replaced recurrence

5/5

A recurrent network reads one position after another, and that sequential nature precludes parallelization within a training example 1. A self-attention layer connects all positions in a constant number of sequential steps, where a recurrent layer needs O(n) 2. The 2017 paper dropped recurrence entirely 3; its models trained faster and reached 28.4 BLEU on English-to-German translation 4. With no recurrence, word order comes from sine and cosine waves of different frequencies 5, a cousin of the Fourier transform.

5 of 5 quotes found in their sources
  1. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org sequential nature precludes parallelization within training examples Quote found in the source
  2. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires Quote found in the source
  3. Attention Is All You Need (Vaswani et al., 2017), arXiv abstract arxiv.org dispensing with recurrence and convolutions entirely Quote found in the source
  4. Attention Is All You Need (Vaswani et al., 2017), arXiv abstract arxiv.org Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task Quote found in the source
  5. Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org sine and cosine functions of different frequencies Quote found in the source

claude-opus-5-5for @vizipediav1 ·

1 version