Pages / #machine-learning / #transformers
Attention in a transformer
Self-attention lets every word in a sentence look at every other word and decide how much each one matters, all at once. It is the core of the transformer.
Where each word looksInteractive
- Introduced
- 2017, "Attention Is All You Need"
- Heads in the original model
- 8
- Key size per head
- 64
- Score scaling
- ÷ √64 = ÷ 8
- Steps between any two words
- Constant (recurrence: O(n))
- English-German BLEU (2017)
- 28.4
One word, four steps
In words
Queries, keys and values
6/6
An attention function maps a query and a set of key-value pairs to an output 1. Each word gets a query, a key and a value vector 2, and its score for another word is the dot product of its query with that word's key 3. The scores are divided by the square root of the key size, then a softmax turns them into weights 4 that are positive and add up to 1 5. The output is the weighted sum of the values 6.
6 of 6 quotes found in their sources
-
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
An attention function can be described as mapping a query and a set of key-value pairs to an output
Quote found in the source -
The Illustrated Transformer (Jay Alammar) jalammar.github.io
So for each word, we create a Query vector, a Key vector, and a Value vector.
Quote found in the source -
The Illustrated Transformer (Jay Alammar) jalammar.github.io
The score is calculated by taking the dot product of the query vector with the key vector of the respective word
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
and apply a softmax function to obtain the weights on the values
Quote found in the source -
The Illustrated Transformer (Jay Alammar) jalammar.github.io
Softmax normalizes the scores so they’re all positive and add up to 1.
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
The output is computed as a weighted sum of the values
Quote found in the source
Many heads at once
4/4
The original Transformer runs 8 attention heads side by side, each with its own query, key and value weights 12. Several heads let the model attend to different kinds of information at different positions 3. In Jay Alammar's walkthrough, while encoding "it", one head focuses most on "the animal" and another on "tired" 4, the two patterns the heads above imitate.
4 of 4 quotes found in their sources
-
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
parallel attention layers, or heads
Quote found in the source -
The Illustrated Transformer (Jay Alammar) jalammar.github.io
eight attention heads, so we end up with eight sets for each encoder/decoder
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions.
Quote found in the source -
The Illustrated Transformer (Jay Alammar) jalammar.github.io
one attention head is focusing most on "the animal", while another is focusing on "tired"
Quote found in the source
Why it replaced recurrence
5/5
A recurrent network reads one position after another, and that sequential nature precludes parallelization within a training example 1. A self-attention layer connects all positions in a constant number of sequential steps, where a recurrent layer needs O(n) 2. The 2017 paper dropped recurrence entirely 3; its models trained faster and reached 28.4 BLEU on English-to-German translation 4. With no recurrence, word order comes from sine and cosine waves of different frequencies 5, a cousin of the Fourier transform.
5 of 5 quotes found in their sources
-
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
sequential nature precludes parallelization within training examples
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), arXiv abstract arxiv.org
dispensing with recurrence and convolutions entirely
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), arXiv abstract arxiv.org
Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task
Quote found in the source -
Attention Is All You Need (Vaswani et al., 2017), full text on arXiv arxiv.org
sine and cosine functions of different frequencies
Quote found in the source