Why a Transformer Cannot See Word Order
Self-attention handles every token at the same time, so it never sees which token came first. This part works one sentence and its reversal by hand: both give the same output vectors, only in a different order. It then shows why nothing later in the model can use that order. The fix is positional encoding, and the part ends by explaining why the simplest version of it, a counter, does not work.
3.1The problem: self-attention cannot see word order
In Part 1 and Part 2 you built the self-attention mechanism. It takes a sequence of tokens, computes queries, keys, and values, and hands back a new sequence. Every output token is a weighted blend of the input tokens.
But the design has a fatal flaw: self-attention doesn't know what order the tokens arrived in.
Compare it with a recurrent neural network, which reads a sentence one word at a time. It knows
dog came before bites because it processed dog first. Self-attention
processes every token at the exact same moment, and it compares every token to every other token with one
matrix multiplication.
Think for a moment If you take the sentencedog bites manand change it toman bites dog, the meaning changes completely. How does the self-attention formula \(\text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d_k}}\right)\mathbf{V}\) react to that change?
It doesn't react at all. The formula does the exact same math either way, because it treats your sentence as a bag of words rather than an ordered sequence. Without a sense of order, the model can't tell who is biting whom.
Figure 1 puts the two side by side. Now let's prove it by running the numbers.
3.2Reverse the sentence, and the same vectors come back
Take the three-word sentence from Part 1, I love LLMs, and push it through self-attention by
hand. Then swap the words to LLMs love I and watch what happens to the output.
The original sequence: I love LLMs
We reuse the setup from Part 1 exactly, so nothing here has to be taken on trust. The
embedding step turns I into \((0, 0, 1, 1)\), love into
\((0, 1, 0, 1)\), and LLMs into \((1, 0, 0, 1)\). Four numbers each, so the
input matrix \(\mathbf{X}\) is 3 × 4. Row 1 is I,
row 2 is love, row 3 is LLMs:
Two words we will keep straight There are two different "wheres" in this part, and they are easy to confuse, so here are the names for them — used this way from here to the end.Same three words either way, so whenever you meet a number in this part, the first question to ask is which of the two it is counting.
- Position is where a word sits in the sentence, counting from 0. In
I love LLMs,Iis at position 0,loveat position 1,LLMsat position 2. Position picks a row of \(\mathbf{X}\).- Slot is where a number sits inside one vector, also counting from 0. In this part that vector is one word's row of \(\mathbf{X}\):
Iis \((0, 0, 1, 1)\), so its slot 0 holds 0, slot 2 holds 1, and slot 3 holds 1. Slot picks a column of \(\mathbf{X}\). Later parts use the word for any vector of numbers, not only for a word's row, and the counting is the same in all of them.One warning about numbering. Two counting systems are about to sit next to each other. Positions and slots count from 0, as above. Rows and columns of a matrix count from 1, which is what everyone writing about matrices does. They point at the same places: row 1 of \(\mathbf{X}\) is the word at position 0, and column 1 holds slot 0 of every row. Outside this box the two counts never share a sentence, so every number you meet next to the word row or column starts at 1, and every number next to position or slot starts at 0.
The three weight matrices are the ones training produced in Part 1, each 4 × 4:
\[ \begin{aligned} \mathbf{W}^Q&=\begin{pmatrix}0&-1&1&1\\0&1&0&0\\1&0&0&-1\\1&1&0&1\end{pmatrix}\\[8pt] \mathbf{W}^K&=\begin{pmatrix}0&-1&0&1\\-1&1&0&0\\1&-1&1&-1\\1&1&0&1\end{pmatrix}\\[8pt] \mathbf{W}^V&=\begin{pmatrix}0&0&1&0\\1&0&0&1\\0&1&0&0\\1&0&0&0\end{pmatrix} \end{aligned} \]Multiplying \(\mathbf{X}\) by each of them gives the Query, Key and Value matrices. Part 1 worked one of those dot products in full and left the pattern for the other thirty-five, so here are the results:
\[ \begin{aligned} \mathbf{Q}=\mathbf{X}\mathbf{W}^Q&=\begin{pmatrix}2&1&0&0\\1&2&0&1\\1&0&1&2\end{pmatrix}\\[8pt] \mathbf{K}=\mathbf{X}\mathbf{W}^K&=\begin{pmatrix}2&0&1&0\\0&2&0&1\\1&0&0&2\end{pmatrix}\\[8pt] \mathbf{V}=\mathbf{X}\mathbf{W}^V&=\begin{pmatrix}1&1&0&0\\2&0&0&1\\1&0&1&0\end{pmatrix} \end{aligned} \]Now compute the attention scores, \(\mathbf{Q}\mathbf{K}^\top\):
\[ \mathbf{Q}\mathbf{K}^\top = \begin{pmatrix} 4 & 2 & 2 \\ 2 & 5 & 3 \\ 3 & 2 & 5 \end{pmatrix} \]Every entry is one dot product between a row of \(\mathbf{Q}\) and a row of \(\mathbf{K}\).
Entry \((1,2)\) asks how well I's query matches love's key:
\((2,1,0,0)\cdot(0,2,0,1) = 0 + 2 + 0 + 0 = 2\). The other eight follow the same way.
Notice this matrix is not symmetric — entry \((1,3)\) is 2 while entry \((3,1)\) is 3. Asking is not the same as being asked, and that asymmetry survives everything below.
Next, divide by \(\sqrt{d_k}\). Here \(d_k = 4\), so that's \(\sqrt{4} = 2\). Then apply softmax to each row to get the weight matrix \(\mathbf{A}\):
\[ \frac{\mathbf{Q}\mathbf{K}^\top}{2} = \begin{pmatrix} 2.0 & 1.0 & 1.0 \\ 1.0 & 2.5 & 1.5 \\ 1.5 & 1.0 & 2.5 \end{pmatrix} \] \[ \mathbf{A} = \operatorname{softmax}\begin{pmatrix} 2.0 & 1.0 & 1.0 \\ 1.0 & 2.5 & 1.5 \\ 1.5 & 1.0 & 2.5 \end{pmatrix} = \begin{pmatrix} 0.58 & 0.21 & 0.21 \\ 0.14 & 0.63 & 0.23 \\ 0.23 & 0.14 & 0.63 \end{pmatrix} \]Everything from here on is rounded to two decimals, so a row may add up to 0.99 rather than 1.
Finally, multiply the weights by the values \(\mathbf{V}\) to get the output \(\mathbf{Z}\):
\[ \mathbf{Z} = \mathbf{A}\mathbf{V} = \begin{pmatrix} 0.58 & 0.21 & 0.21 \\ 0.14 & 0.63 & 0.23 \\ 0.23 & 0.14 & 0.63 \end{pmatrix} \begin{pmatrix}1&1&0&0\\2&0&0&1\\1&0&1&0\end{pmatrix} = \begin{pmatrix} 1.21 & 0.58 & 0.21 & 0.21 \\ 1.63 & 0.14 & 0.23 & 0.63 \\ 1.14 & 0.23 & 0.63 & 0.14 \end{pmatrix} \]The first row is the output for I: \((1.21, 0.58, 0.21, 0.21)\).
The second row is the output for love: \((1.63, 0.14, 0.23, 0.63)\).
The third row is the output for LLMs: \((1.14, 0.23, 0.63, 0.14)\).
The reversed sequence: LLMs love I
Now feed the model the reversed sentence, LLMs love I. Row 1 is now LLMs, row
2 is love, and row 3 is I.
The weight matrices haven't changed, so \(\mathbf{Q}\), \(\mathbf{K}\), and \(\mathbf{V}\) just swap their rows too:
Why do the weight matrices not move too?
Hold on, you might say: we shuffled the rows of \(\mathbf{X}\), so why not shuffle \(\mathbf{W}^Q\), \(\mathbf{W}^K\) and \(\mathbf{W}^V\) as well?
They are a different kind of object. \(\mathbf{X}\) is data — it is
this sentence, and it changes the moment you type a different one.
\(\mathbf{W}^Q\) is a parameter. Training set its numbers once, and they
stay put for every sentence the model will ever see, in every language, of every
length. Nothing inside \(\mathbf{W}^Q\) could shuffle, because it was never tied to
your sentence in the first place. Figure 2 draws this with the real numbers:
I, love and LLMs each go through the same
\(\mathbf{W}^Q\), and each comes out as its own row of \(\mathbf{Q}\).
Every row is one of the rows you already computed, sitting somewhere new. No new arithmetic happened.
Now compute the new scores, \(\mathbf{Q}_{\text{swapped}}\mathbf{K}_{\text{swapped}}^\top\):
\[ \mathbf{Q}_{\text{swapped}}\mathbf{K}_{\text{swapped}}^\top = \begin{pmatrix} 5 & 2 & 3 \\ 3 & 5 & 2 \\ 2 & 2 & 4 \end{pmatrix} \]Think for a moment Put the two score matrices side by side. They aren't the same matrix. Is anything about them the same?
Look at the two sets of nine numbers:
\[ \mathbf{Q}\mathbf{K}^\top=\begin{pmatrix} 4 & 2 & 2 \\ 2 & 5 & 3 \\ 3 & 2 & 5 \end{pmatrix} \qquad \mathbf{Q}_{\text{swapped}}\mathbf{K}_{\text{swapped}}^\top=\begin{pmatrix} 5 & 2 & 3 \\ 3 & 5 & 2 \\ 2 & 2 & 4 \end{pmatrix} \]The same nine numbers are present in both, and the second matrix is the first one flipped: the row order reversed, and the column order reversed with it. The 4 that sat in the top-left corner is now in the bottom-right corner.
Why it had to happen: X appears twice in the score
Write the score matrix in terms of \(\mathbf{X}\):
\[ \mathbf{Q}\mathbf{K}^\top = \mathbf{X}\mathbf{W}^Q\,(\mathbf{X}\mathbf{W}^K)^\top = \mathbf{X}\,\mathbf{W}^Q\mathbf{W}^{K\top}\,\mathbf{X}^\top \]\(\mathbf{X}\) appears twice. On the left it supplies the rows of the result, and on
the right, as \(\mathbf{X}^\top\), it supplies the columns. Reverse the rows of
\(\mathbf{X}\) and both copies change, so the rows of the score matrix reverse and its
columns reverse with them. No score changes value, because each one is a single
word's row of \(\mathbf{X}\), times \(\mathbf{W}^Q\mathbf{W}^{K\top}\), times another
word's row. For example, I asking and love answering is
in either sentence. That depends on which two words they are, and not on where they sit.
Figure 3 follows two scores. The 4 is I asking and I
answering, so it moves from the top-left corner to the bottom-right corner. The 2 is
I asking and love answering. Its row moves from the top to
the bottom, but its column stays, because love is in the middle of both
sentences. Appendix A.1 shows the same thing for any reordering,
not only a reversal.
Softmax and the output
Everything downstream inherits this. Softmax works on one row at a time. It never looks at the row number, and inside a row it turns each score into its share of the row's total, which is the same in whatever order the scores are added. So the attention weights \(\mathbf{A}\) go through the same change — the same numbers, moved the same way:
\[ \mathbf{A}=\begin{pmatrix} 0.58 & 0.21 & 0.21 \\ 0.14 & 0.63 & 0.23 \\ 0.23 & 0.14 & 0.63 \end{pmatrix} \qquad \mathbf{A}_{\text{swapped}} = \begin{pmatrix} 0.63 & 0.14 & 0.23 \\ 0.23 & 0.63 & 0.14 \\ 0.21 & 0.21 & 0.58 \end{pmatrix} \]Before going on to the output, read one row of these. love is the
middle word, and it stays in the middle when the sentence is reversed. Row 2 of
\(\mathbf{A}\) says how much love paid attention to each word. Here it
is in both sentences, with the words written above the numbers so you can see who is
who:
In the first sentence, love gives I a weight of 0.14, and
I is the word on its left. In the second sentence,
love gives I a weight of 0.14, and I is now the
word on its right. Same word, opposite side, identical weight. And
LLMs gets 0.23 from either side too.
"Attention did not use the order" means this, in numbers: every weight depends on which word and never on which side.
Now multiply the weights by the new values \(\mathbf{V}_{\text{swapped}}\) to get the final output:
\[ \mathbf{Z}_{\text{swapped}} = \mathbf{A}_{\text{swapped}}\mathbf{V}_{\text{swapped}} = \begin{pmatrix} 0.63 & 0.14 & 0.23 \\ 0.23 & 0.63 & 0.14 \\ 0.21 & 0.21 & 0.58 \end{pmatrix} \begin{pmatrix}1&0&1&0\\2&0&0&1\\1&1&0&0\end{pmatrix} = \begin{pmatrix} 1.14 & 0.23 & 0.63 & 0.14 \\ 1.63 & 0.14 & 0.23 & 0.63 \\ 1.21 & 0.58 & 0.21 & 0.21 \end{pmatrix} \]Look closely at \(\mathbf{Z}_{\text{swapped}}\).
Row 1 is the output for LLMs: \((1.14, 0.23, 0.63, 0.14)\).
Row 2 is the output for love: \((1.63, 0.14, 0.23, 0.63)\).
Row 3 is the output for I: \((1.21, 0.58, 0.21, 0.21)\).
Those are the exact same vectors you got the first time. They just changed rows. The vector produced for
I is exactly the same whether I came first or last. The self-attention layer
completely ignored your word order — Figure 4 is the whole section in one picture.
I,
love, and LLMs unchanged. The model cannot tell which word came first.
Does this hold in general?
You have seen one case, and you might reasonably blame the result on the setup: one particular sentence and one particular set of weights. The setup isn't the reason. The same result holds for every embedding, every weight matrix and every sentence length, and it has a name: self-attention is permutation-equivariant — shuffle the input and you get the original output, shuffled the same way.
The proof takes about half a page and needs one new idea (writing a shuffle as a matrix). Read the proof whenever you like, but not right now: nothing in the rest of this part depends on it. It is Appendix A.1.
Key takeaway
- Position is where a word sits in the sentence, counting from 0. Slot is where a number sits inside one vector, also counting from 0. Rows and columns of a matrix count from 1, and the two counts never share a sentence outside the box above.
- Shuffle the rows of \(\mathbf{X}\) and every output vector comes back unchanged,
in the shuffled order. Worked here for
I love LLMsagainstLLMs love I. - The reason is that every step — the three projections, the dot products, the softmax, the weighted sum — reads a row's contents and never its address.
Try it (3 minutes): Swap the first and third rows of \(\mathbf{X}\) and redo \(\mathbf{Q}\mathbf{K}^\top\) by hand. You should find that every entry of the new matrix is an entry of the old one, sitting at a different address. Then say which single entry did not move at all, and why.
3.3Why the same vector is the wrong answer
Section 3.2 left you with a fact: shuffle the words and the same output vectors come back, just in different rows. This section and the next are about what that means for the model.
You probably have three objections at this point, and all three are fair:
- Is the identical output really a problem?
- The rows are still in order, so is the order still there?
- Couldn't a later layer just read the row numbers and use them?
This section answers the first. Section 3.4 answers the other two.
Wait — should it not be the same?
You might be thinking: in I love LLMs and LLMs love I,
I is the same word. So it seems right that the layer gives I
the same output vector in both sentences. Why would a different vector be better?
That would be right if the layer's job were to describe the word on its own, but
that is not its job. The embedding already does that: it gives bank one
fixed vector, the same in every sentence. The attention layer exists to describe the
word as it is used in this sentence. The bank in "river bank" is
a different thing from the bank in "bank account", and the embedding
cannot tell them apart. Attention can: its output vector for bank changes
when the words around bank change.
Each word's output vector is its row of \(\mathbf{Z}\). Ideally, that vector would depend on two things:
- which words are around it
- where those words are
Attention handles the first, as the
bank example shows. Section 3.2 showed that it does not handle the
second: moving the words around left every output vector unchanged.
dog bites man, man bites dog
Take two sentences built from the same three words:
\[ \texttt{dog}\ \texttt{bites}\ \texttt{man} \qquad\text{against}\qquad \texttt{man}\ \texttt{bites}\ \texttt{dog} \]One describes an ordinary afternoon. The other would be on the front page. Run both through the layer. The bag of words is identical, so by the argument above every output vector is identical too — the three vectors simply come out in a different order.
Look at the middle word. bites sits at position 1 in both sentences, so
the reordering doesn't even move it. Its output vector is the same vector, down to
the last decimal, in the ordinary sentence and in the front-page one. And
bites is exactly the word whose meaning depends on the order — it's
the word that has to know who is doing the biting. It comes out of the layer with no
opinion on the matter.
So the identical output is not a sign that the layer is consistent. It shows that the word order never reached the contents of the rows. The rest of this part is about putting the order into the rows.
Key takeaway
- The output of an attention layer is meant to describe a word as it appears in
this sentence. Getting the same vector for
bitesindog bites manandman bites dogis the failure, not the feature. - The identical output does not mean the layer is consistent. It means the word order never reached the contents of the rows.
Try it (2 minutes): Write dog bites man and
man bites dog one under the other. For each of the three words, say whether
it moved, and whether its output vector changes. You should get two words that moved
and zero vectors that changed. Then say in one sentence why that count would be zero
for any reordering of these three words, not only this swap.
3.4The order is in the row numbers, and nothing reads them
Think for a moment After the swap, every output vector stayed the same, and the three vectors came out in a new order that follows the words. So is the word order lost, or is it still somewhere in the output?
Yes, the order is still there
Yes, the order is still there. Row 1 of \(\mathbf{Z}\) is still the first word and row 3 is still the last. Anyone who looks at the row numbers of \(\mathbf{Z}\) can see the order. Attention did not use them: that is what Sections 3.2 and 3.3 showed.
The row numbers really are still there for anyone to see, and this was attention's only chance to use them — the layer is finished, and the next thing these vectors meet is the feed-forward network. So it is fair to ask whether something further down the model could pick the row numbers up instead. That is the rest of this section.
Nothing after attention reads it either
First, get precise about what "reading the row number" would take. A component uses position only if some parameter of it is different for different positions — a weight matrix \(\mathbf{W}^{(0)}\) reserved for position 0, a different \(\mathbf{W}^{(1)}\) for position 1, and so on. Then position 0 and position 2 get different treatment, and the model can tell them apart.
If no step has weights like that, the row number is only a label, and nothing in the model reacts to it. So the question is: does any step after attention use one weight matrix for position 0, a different one for position 1, and so on?
The feed-forward network doesn't. The feed-forward network is the step that sits after attention inside every Transformer layer. Part 5 builds it properly, and all we need here is its shape. It never looks at the sequence at all. It takes one row at a time and pushes it through a small two-layer network:
\[ \operatorname{FFN}(\mathbf{z}_i)=\max(0,\ \mathbf{z}_i\mathbf{W}_1+\mathbf{b}_1)\,\mathbf{W}_2+\mathbf{b}_2 \]Read the subscripts. The row index \(i\) appears on \(\mathbf{z}_i\) and nowhere else. \(\mathbf{W}_1, \mathbf{b}_1, \mathbf{W}_2, \mathbf{b}_2\) carry no \(i\) — there's one set of them, shared by every position in every sentence. Computing row 1 and computing row 3 is literally the same arithmetic with different numbers fed in.
Figure 2 already drew this situation. One shared set of weights, applied to each row on its own, with the row number used as an address and never as an ingredient. Shuffle the rows going in and the rows come out shuffled, contents untouched.
The next attention layer doesn't either. The proof in Appendix A.1 never mentions which layer it's talking about. Layer two is exactly as order-blind as layer one, for exactly the same reason.
Put those together. Every block in the stack passes a shuffle straight through, and blocks that pass a shuffle through still pass it through when you stack them. So the whole network behaves like one big order-blind layer: shuffle the words and every token's final vector is unchanged, just relocated. The row numbers were available the entire way down and not one parameter was wired to notice them.
Why no component gets its own weights per position
All of this could be avoided by giving each position its own weights: \(\mathbf{W}^{(0)}\) for position 0, \(\mathbf{W}^{(1)}\) for position 1, out to \(\mathbf{W}^{(499)}\). Nobody does, and it is worth seeing why, because it explains why position has to arrive some other way.
- The length gets frozen. Train with 500 positions and a 501st token has no weights to be processed by. The model simply cannot read a longer sentence. GPT-2 has this limit for a related reason. GPT-2 uses a learned position vector for each of 1024 positions, so it cannot read a longer text.
- The later positions barely train. \(\mathbf{W}^{(437)}\) only receives a gradient from sentences of 438 tokens or more. Sentences that long are rare, so that matrix stays close to whatever it was initialised as.
- Every rule has to be learned 500 times. "An adjective usually comes before its noun" would be learned at position 5, and position 300 would gain nothing from that lesson and would have to discover the same rule again from scratch.
Sharing one set of weights across all positions buys the opposite of all three: any length, dense training signal, and one rule that works everywhere. You would be giving up far too much. So the order-blindness isn't an oversight — it is the price of sharing one set of weights. So position cannot come from giving each position its own weights.
After the encoder, the row numbers are gone
In a translation model, the layers that read the English sentence are the encoder. The part that writes the translation is the decoder. The decoder reads the encoder's output through cross-attention, which Part 1, Section 1.2.5 introduced and Part 7 works through with numbers. Cross-attention is still attention, so what it hands the decoder is a weighted sum of the encoder's rows. A sum comes out the same in whatever order its terms are added. Reverse the English sentence: every row keeps its numbers and every weight stays with its row, so the sum is the same vector. Appendix A.2 writes this out step by step, and it is meant to be read after Part 7. Figure 5 shows where cross-attention sits between the two.
Cross-attention is where the order is really lost. Inside the encoder, every output has
one row per English word, so the order at least survives as the row numbers. The output
of cross-attention has one row per word the decoder writes. The English rows have been
added into one vector, and there is nowhere left to keep an English row number. So the
decoder receives exactly the same thing for dog bites man and
man bites dog, and it cannot translate them differently.
A row number cannot survive a sum, but the contents of a row can. That is why the position has to change the contents of the rows before any step adds rows together.
What is left: put the position inside the vector
Think for a moment Everything the model computes comes from two kinds of numbers: the weights, which are the same for every position, and the contents of the rows. This section has ruled out the weights. Where does that leave the position?
In the rows. Suppose the model is to treat I at position 0 differently
from I at position 2. Then the contents of I's row have to be
different in the two sentences. Right now they are not: I's row starts
as its embedding, \((0, 0, 1, 1)\), in both.
So something that depends on the position has to be added to that row before
attention runs. I arrives as \((0, 0, 1, 1)\) plus a vector that means
"position 0" in the first sentence, and as \((0, 0, 1, 1)\) plus a vector that
means "position 2" in the second. That is positional encoding in one sentence: add to each word's
embedding a vector that says where the word sits. From then on the position is
part of what every dot product reads, so it reaches every attention step and survives
every sum. Section 3.5 builds exactly that. It is not the only way: Part 4 shows RoPE,
which leaves the embedding alone and changes the query and the key inside attention
instead.
✗ Common mistake The row numbers are right there, so it is tempting to imagine that a component bolted onto the end, one that reads position 0 directly, would be enough. That is wrong. Reading position 0 answers one question: which word is sitting there. "Alice told Bob about Carol" and "Alice told Carol about Bob" both answerAlice, andAliceleaves the stack as the same vector in both, so anything the model says about her is the same in both.
Key takeaway
- The order survives only as row numbers, and no step that computes a row's contents ever reads a row number. The one step that does read row numbers is the decoder's mask, which only decides which rows a row may see. Part 8 builds it.
- A step could read row numbers only if it had different weights for different positions. No step does, because that would fix the sentence length and make every rule be learned once per position.
- Cross-attention adds the encoder's rows into one vector. After that, even the row numbers are gone.
- So the position has to be inside the rows. Positional encoding puts it there by adding a position vector to each word's embedding (Section 3.5).
3.5The fix: positional encoding
If the model can't see the order of the tokens, you have to write the order into the tokens themselves.
A positional encoding is a vector of numbers added to a token to tell the network where that token sits in the sequence.
Think of a ticket queue at a bakery. The counter staff process whoever comes up to the window. If you want them to know who arrived first, you have to hand every customer a numbered ticket.
Before the embedded tokens \(\mathbf{X}\) reach the attention layer, add a second matrix to them. That second matrix holds the position tickets: one row per position, the same width as a row of \(\mathbf{X}\). Call it \(\mathbf{P}\).
\[ \mathbf{X}_{\text{input}} = \mathbf{X} + \mathbf{P} \]That sum gets a name, because from here on it's not the embedding any more. The embedding says what a word means; the sum says what it means and where it sits. It's called the input vector, since it's what actually enters the first attention layer. Everything downstream — \(\mathbf{Q}\), \(\mathbf{K}\), \(\mathbf{V}\), and all the arithmetic of Sections 3.2 to 3.4 — is built from input vectors, not from embeddings.
If a token is at position 0, it gets the position 0 encoding added to it. If it moves to position 2, it gets
the position 2 encoding added to it instead. Now, if you swap the sequence to LLMs love I, the token
I gets a different positional encoding. Its input numbers change, so its output numbers change. The
model can finally see the order.
Key takeaway
- A positional encoding is a vector that says where a word sits. It is added to the word's embedding before the first attention layer.
- The sum is the input vector: what the word means plus where it sits. From here on, \(\mathbf{Q}\), \(\mathbf{K}\), \(\mathbf{V}\) and all the attention arithmetic are computed from input vectors, not from embeddings.
- Move a word and it gets a different positional encoding, so its input vector changes and its output changes. That is how the model finally sees the order.
3.6The same example with positions added: different vectors now
Watch what happens to the word I once you add positional encodings to the sequence. To keep it
simple, say position 0 adds a vector \(\mathbf{p}_0\), position 1 adds \(\mathbf{p}_1\), and position 2 adds
\(\mathbf{p}_2\).
In the original sequence I love LLMs, the word I is at position 0. So its input
vector becomes \(\mathbf{x}_{\text{I}} + \mathbf{p}_0\). This combined vector is multiplied by the weight matrix
\(\mathbf{W}^Q\) to create its query vector:
Now take the swapped sequence, LLMs love I. The word I is now at position 2. Its
input vector becomes \(\mathbf{x}_{\text{I}} + \mathbf{p}_2\). Its query vector becomes:
\(\mathbf{p}_0\) and \(\mathbf{p}_2\) are completely different vectors, so I now generates a
completely different query vector. It asks different questions of the rest of the sentence, produces different
attention scores, and ends up with a completely different output vector \(\mathbf{z}\).
I becomes a different
input vector depending on where it sits in the sentence. This causes it to generate a different Query vector,
produce different attention scores, and result in a different final output.Figure 6 traces that through. The self-attention layer is still doing the same math, but the input numbers have changed, so the output
numbers change too, and the model can finally tell I love LLMs apart from
LLMs love I.
3.7Why not use a simple counter?
You might be thinking: why do we need whole vectors for this? Why not just use a counter — add 1 to everything at position 1, 2 to everything at position 2, and so on?
What a position system has to make possible
Before judging the counter, say what any position system has to make possible.
Words relate to each other by relative position: the adjective before the noun, the verb two words back. So the model must be able to learn rules like "look at the word three steps back". Say the query's word is at position \(m\) and the key's word is at position \(n\). Such a rule cares about \(m - n\), how far apart the two words are. It does not care about \(m\) and \(n\) themselves, where in the sentence they sit. The gap is 3 for positions 4 and 1 and also for positions 400 and 397, and the rule has to work in both places.
Now turn that into a statement about scores. Position can only enter the model through the dot product between a query and a key (Section 3.4). A high score means "attend to this word". So "look three steps back" means two things. The score between positions \(m\) and \(n\) has to be highest when \(m - n = 3\). And it has to be the same score for every \(m\).
The requirement The score between the query at position \(m\) and the key at position \(n\) must have a part that can depend on \(m - n\) alone, and not on \(m\) or \(n\) themselves. Then two pairs with the same gap can get the same position score, wherever they sit in the sentence.
Two things wrong with a counter
The first is easy to see. In a sentence with 500 words, the last token would have 499 added to its values. That number would drown out the embedding, and the model would see the position instead of the word.
The second is the important one: a counter cannot meet the requirement. Showing why takes the rest of this section.
Where in the score does position actually live?
One thing to settle before any numbers. The vector entering attention is the input vector, embedding plus encoding. Write \(\mathbf{e}_m\) for the embedding vector of whatever word sits at position \(m\) — it is the \(\mathbf{x}\) of Section 3.6, now carrying a position label — so the query and key are \(\mathbf{q}_m = (\mathbf{e}_m + \mathbf{p}_m)\mathbf{W}^Q\) and \(\mathbf{k}_n = (\mathbf{e}_n + \mathbf{p}_n)\mathbf{W}^K\). The embedding certainly isn't zero. Multiply those out and the score isn't one thing, it's four:
\[ (\mathbf{e}_m + \mathbf{p}_m)\,\mathbf{M}\,(\mathbf{e}_n + \mathbf{p}_n)^\top = \underbrace{\mathbf{e}_m \mathbf{M} \mathbf{e}_n^\top}_{\text{word, word}} + \underbrace{\mathbf{e}_m \mathbf{M} \mathbf{p}_n^\top}_{\text{word, position}} + \underbrace{\mathbf{p}_m \mathbf{M} \mathbf{e}_n^\top}_{\text{position, word}} + \underbrace{\mathbf{p}_m \mathbf{M} \mathbf{p}_n^\top}_{\text{position, position}} \]Here \(\mathbf{M} = \mathbf{W}^Q \mathbf{W}^{K\top}\) is whatever the two weight matrices multiply out to — the model learns it. (You'll meet this expansion again in Part 4, where it's the reason RoPE exists.)
The requirement asks for a part of the score that depends on \(m - n\). Which of the four terms could that be? The positions enter the score only through the two position vectors, \(\mathbf{p}_m\) and \(\mathbf{p}_n\). A term can depend on \(m - n\) only if it contains both, so look for the two position vectors in each term:
| term | contains \(\mathbf{p}_m\), the query's position? | contains \(\mathbf{p}_n\), the key's position? |
|---|---|---|
| \(\mathbf{e}_m \mathbf{M} \mathbf{e}_n^\top\) | no | no |
| \(\mathbf{e}_m \mathbf{M} \mathbf{p}_n^\top\) | no | yes |
| \(\mathbf{p}_m \mathbf{M} \mathbf{e}_n^\top\) | yes | no |
| \(\mathbf{p}_m \mathbf{M} \mathbf{p}_n^\top\) | yes | yes |
The second and third terms each carry one position only: the second has \(\mathbf{p}_n\) but not \(\mathbf{p}_m\), and the third has \(\mathbf{p}_m\) but not \(\mathbf{p}_n\). Such a term is not useless, because it lets a word care about an absolute position, such as the start of the sentence. But \(m - n\) needs both positions, and only the fourth term contains both position vectors.
So the demonstration below looks only at \(\mathbf{p}_m \mathbf{M} \mathbf{p}_n^\top\). The other three terms are not zero, but none of them can carry the rule.
The counter, worked through
The counter proposal was to add the position number to everything, so with the 4-number vectors of this example the encoding for position \(m\) is \(\mathbf{p}_m = (m,\ m,\ m,\ m)\). Write that as \(m\) times the all-ones vector, \(\mathbf{p}_m = m\mathbf{1}\), and the fourth term collapses:
\[ \mathbf{p}_m \mathbf{M} \mathbf{p}_n^\top = (m\mathbf{1})\,\mathbf{M}\,(n\mathbf{1})^\top = mn \cdot \underbrace{\mathbf{1}\mathbf{M}\mathbf{1}^\top}_{\text{some number } c} = c\,mn \]Read what that says. Whatever the model learns for \(\mathbf{W}^Q\) and \(\mathbf{W}^K\), the position part of the score is some constant times \(mn\). The weights get to pick \(c\), and that's all they get to pick. So you don't have to guess at the weights — you can check every possible model at once by looking at \(mn\).
Take three pairs of positions that are all exactly 3 apart:
\[ \begin{array}{llr} \text{positions 3 and 0:} & c \times 3 \times 0 & = 0 \\[4pt] \text{positions 10 and 7:} & c \times 10 \times 7 & = 70c \\[4pt] \text{positions 100 and 97:} & c \times 100 \times 97 & = 9700c \end{array} \]One gap, three completely different scores. The requirement asked for the same position score for every pair that is 3 apart. A counter gives \(0\), \(70c\) and \(9700c\), so it fails the requirement.
There is a second problem. The requirement asked for the word 3 back to get the highest score. The counter does not give it that. The word at position 10 scores \(70c\) against the word at position 7 — the one it's supposed to pick — but \(100c\) against itself, because \(10 \times 10\) beats \(10 \times 7\). Softmax would hand the attention to the wrong word.
These scores measure absolute position, how far into the sentence the pair sits, and not relative position. "Three back" has no score of its own, so there is no rule the model could learn from these scores that means "three back".
✗ Common mistake The numbers above are large, so it is tempting to read this as reason one again, fixable by scaling the counter down, but that is wrong. Scaling only changes \(c\), and \(c\) multiplies all four scores equally — \(0\), \(70c\), \(9700c\) and \(100c\) keep exactly the same ranking however small you make it. The trouble is the shape of \(mn\), not how big it gets.
Part 4 runs these same three pairs of positions through the sine and cosine encoding.
Key takeaway
- A position system has to let the model learn rules about relative position, such as "look three words back". For that, some part of the score must be able to depend on \(m - n\) alone.
- A counter cannot do this. Its position score is \(c\,mn\), so pairs that are 3 apart score \(0\), \(70c\) and \(9700c\). There is no single value that means "3 apart".
Try it (3 minutes): Find another pair of positions three apart whose counter score is \(1600c\). Then find the pair three apart with the smallest possible score. Having done both, say in one sentence why "three apart" cannot be read off a counter score.
AAppendix: the proof in full
This was pulled out of the main text so that a first reading runs without stopping. Nothing later depends on it. Read it when you want the general case rather than the worked example, or when you do not believe the worked example.
A.1Self-attention is permutation-equivariant
This one belongs to Section 3.2. There we ran one sentence and its reversal through attention by hand, and got the same three output vectors in a different order. Here is the same claim for any embeddings, any weight matrices and any sentence length.
Shuffling the rows of a matrix is itself a matrix multiplication. Let \(\mathbf{P}\) be a permutation matrix — the identity with its rows put in a different order. Multiplying on the left by \(\mathbf{P}\) reorders rows, so \(\mathbf{P}\mathbf{X}\) is the shuffled sentence. The swap in Section 3.2 used
\[ \mathbf{P}=\begin{pmatrix}0&0&1\\0&1&0\\1&0&0\end{pmatrix} \]One property of \(\mathbf{P}\) does all the work: undoing a shuffle is the same as shuffling back, so \(\mathbf{P}^\top\mathbf{P} = \mathbf{I}\).
Now feed the shuffled input through attention. Since \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) come from \(\mathbf{X}\) by multiplying on the right, shuffling the rows of \(\mathbf{X}\) shuffles their rows too. The scores become
\[ (\mathbf{P}\mathbf{Q})(\mathbf{P}\mathbf{K})^\top =\mathbf{P}\mathbf{Q}\mathbf{K}^\top\mathbf{P}^\top \]which is the original score matrix with its rows and columns shuffled the same way. Softmax runs along each row and divides by that row's total; reordering the entries of a row reorders the results but does not change their values, because the total is the same however you add the entries up. So softmax passes straight through the shuffle:
\[ \operatorname{softmax}\!\left(\frac{\mathbf{P}\mathbf{Q}\mathbf{K}^\top\mathbf{P}^\top}{\sqrt{d_k}}\right) =\mathbf{P}\,\mathbf{A}\,\mathbf{P}^\top \]Finally multiply by the shuffled values, and watch the two \(\mathbf{P}\) terms in the middle cancel:
\[ \mathbf{Z}_{\text{shuffled}} =\mathbf{P}\mathbf{A}\underbrace{\mathbf{P}^\top\mathbf{P}}_{=\,\mathbf{I}}\mathbf{V} =\mathbf{P}\mathbf{A}\mathbf{V} =\mathbf{P}\mathbf{Z} \]Read the ends of that chain: shuffle the input and you get the original output, shuffled the same way. Not a similar output — the same numbers, in a different order. This holds for every embedding, every weight matrix, every sentence length — which is what Section 3.2 called permutation equivariance.
Think for a moment The cancellation happened because \(\mathbf{P}^\top\mathbf{P} = \mathbf{I}\) sits between \(\mathbf{A}\) and \(\mathbf{V}\). Where exactly did the position information go — was it destroyed, or was it never there?
It was never there. Nothing in \(\mathbf{Q}\mathbf{K}^\top\) refers to row numbers; the entry in row \(i\), column \(j\) is a dot product of two vectors, and a dot product does not know where its operands were sitting. Attention is not losing order — it never had it. So the fix has to come from outside the mechanism, by writing position into the vectors themselves.
A.2The row numbers do not reach the decoder
This one belongs to Section 3.4, and it is written for a second reading. It uses cross-attention as Part 7 builds it, with Part 7's names for the matrices. If you have not read Part 7 yet, read Part 7, Sections 7.2 and 7.4 first and then come back. Nothing in the rest of this part depends on this appendix.
Section 3.4 stated that without a positional encoding, an English sentence and its
reversal give the decoder exactly the same input. Here is why, for
I love LLMs and LLMs love I, one step at a time.
The names
\(\mathbf{C}\) is the encoder's output for I love LLMs: three rows,
one per English word, in sentence order. Call the rows
\(\mathbf{c}_{\texttt{I}}\), \(\mathbf{c}_{\texttt{love}}\) and
\(\mathbf{c}_{\texttt{LLMs}}\). \(\mathbf{P}\) is the permutation matrix of
Appendix A.1, the one that reverses three rows, so
\(\mathbf{q}\) is one row of \(\mathbf{Q}_{\text{cross}}\): the Query of one decoder row. It is built from the decoder's own rows, so it is the same whichever English sentence the encoder read. \(\mathbf{W}^{K}_{\text{cross}}\) and \(\mathbf{W}^{V}_{\text{cross}}\) are the two cross-attention weight matrices that the encoder's output is multiplied by.
Step 0: the encoder's output is reordered, and nothing else
Section 3.4 argued that every step inside an encoder block passes a shuffle straight through. Attention does, by Appendix A.1. The feed-forward network does, because it works on one row at a time, and so do the Add and Norm of Part 5. Six blocks in a row still pass it through. So feeding the encoder the reversed sentence gives
\[ \text{Encoder}(\mathbf{P}\mathbf{X})=\mathbf{P}\,\text{Encoder}(\mathbf{X})=\mathbf{P}\mathbf{C}. \]The three rows of \(\mathbf{C}\) come out in the reverse order, and each row keeps its numbers.
Step 1: the Keys and the Values are reordered, and nothing else
For the original sentence, \(\mathbf{K}=\mathbf{C}\mathbf{W}^{K}_{\text{cross}}\) and \(\mathbf{V}=\mathbf{C}\mathbf{W}^{V}_{\text{cross}}\). For the reversed sentence,
\[ \mathbf{K}'=(\mathbf{P}\mathbf{C})\mathbf{W}^{K}_{\text{cross}} =\mathbf{P}(\mathbf{C}\mathbf{W}^{K}_{\text{cross}})=\mathbf{P}\mathbf{K}, \qquad \mathbf{V}'=\mathbf{P}\mathbf{V}. \]Matrix multiplication can be grouped either way. Reordering the rows and then multiplying by \(\mathbf{W}^{K}_{\text{cross}}\) gives the same result as multiplying first and reordering after. Written as rows:
\[ \mathbf{K}=\begin{pmatrix}\mathbf{k}_{\texttt{I}}\\ \mathbf{k}_{\texttt{love}}\\ \mathbf{k}_{\texttt{LLMs}}\end{pmatrix}, \qquad \mathbf{K}'=\begin{pmatrix}\mathbf{k}_{\texttt{LLMs}}\\ \mathbf{k}_{\texttt{love}}\\ \mathbf{k}_{\texttt{I}}\end{pmatrix}. \]Each word's Key is the same vector in both. Only its row has changed. The same holds for the Values.
Step 2: the scores are reordered, and nothing else
The three scores for the original sentence are one dot product per Key:
\[ \mathbf{s}=\frac{\mathbf{q}\mathbf{K}^\top}{\sqrt{d_k}} =\frac{1}{\sqrt{d_k}}\bigl(\mathbf{q}\cdot\mathbf{k}_{\texttt{I}},\ \mathbf{q}\cdot\mathbf{k}_{\texttt{love}},\ \mathbf{q}\cdot\mathbf{k}_{\texttt{LLMs}}\bigr) =(s_{\texttt{I}},\ s_{\texttt{love}},\ s_{\texttt{LLMs}}). \]For the reversed sentence, use \((\mathbf{P}\mathbf{K})^\top=\mathbf{K}^\top\mathbf{P}^\top\):
\[ \mathbf{s}'=\frac{\mathbf{q}(\mathbf{P}\mathbf{K})^\top}{\sqrt{d_k}} =\frac{\mathbf{q}\mathbf{K}^\top}{\sqrt{d_k}}\,\mathbf{P}^\top =\mathbf{s}\mathbf{P}^\top =(s_{\texttt{LLMs}},\ s_{\texttt{love}},\ s_{\texttt{I}}). \]Multiplying a row vector on the right by \(\mathbf{P}^\top\) reorders its entries the same way \(\mathbf{P}\) reorders rows. Look at one entry. \(\mathbf{q}\cdot\mathbf{k}_{\texttt{I}}\) is a dot product of the same two vectors in both sentences. So it has the same value whether \(\mathbf{k}_{\texttt{I}}\) sits in the first row or the last.
Step 3: the weights are reordered, and nothing else
Softmax turns the three scores into three weights. The weight for I is
The top depends only on \(s_{\texttt{I}}\). The bottom is a sum of three numbers, and
a sum is the same in whatever order its terms are added. So I gets the same
weight in both sentences, and the same goes for the other two words:
Step 4: the weighted sum does not change at all
The output for this decoder row is the weights times the Values. For the original sentence:
\[ \text{output}=\mathbf{a}\mathbf{V} =a_{\texttt{I}}\mathbf{v}_{\texttt{I}}+a_{\texttt{love}}\mathbf{v}_{\texttt{love}}+a_{\texttt{LLMs}}\mathbf{v}_{\texttt{LLMs}}. \]For the reversed sentence, the weights are reordered and the Values are reordered the same way, so the three products pair up exactly as before:
\[ \text{output}'=\mathbf{a}'\mathbf{V}' =a_{\texttt{LLMs}}\mathbf{v}_{\texttt{LLMs}}+a_{\texttt{love}}\mathbf{v}_{\texttt{love}}+a_{\texttt{I}}\mathbf{v}_{\texttt{I}}. \]The same three terms, added in a different order, give the same vector. In matrix form the whole argument is one line, and it is the cancellation of Appendix A.1 again:
\[ \text{output}'=(\mathbf{a}\mathbf{P}^\top)(\mathbf{P}\mathbf{V}) =\mathbf{a}\underbrace{\mathbf{P}^\top\mathbf{P}}_{=\,\mathbf{I}}\mathbf{V} =\mathbf{a}\mathbf{V}=\text{output}. \]Step 5: the same argument covers the real model
Steps 1 to 4 followed one head, one decoder row and one decoder block. A real model has several of each, and the output is still the same for both sentences.
Several heads. Part 2 split attention into heads, and Part 7 says cross-attention is split the same way. Each head has its own \(\mathbf{W}^{K}\) and \(\mathbf{W}^{V}\). So each head runs Steps 1 to 4 on its own, and each head gets the same output for both sentences. The heads' outputs are then placed side by side and multiplied by \(\mathbf{W}^{O}\). Both of those steps take the heads' outputs as they are, and those outputs are already the same for both sentences.
Several decoder rows. From the second run on, the decoder has more than one row, so there is more than one \(\mathbf{q}\). Each row's output is computed from its own \(\mathbf{q}\) and the shared \(\mathbf{K}\) and \(\mathbf{V}\), so Steps 2 to 4 hold for each row separately.
Several decoder blocks. The first block's cross-attention gives the same output for both sentences. The other steps in that block, self-attention and the feed-forward network, read only the decoder's own rows, never the encoder's. So the row that leaves the first block is the same for both sentences. The second block receives the same input, and the same reasoning repeats up to the top of the stack.
Think for a moment Read Steps 0 to 4 again and look for the place where an English row number could have entered the calculation. Is there one?
There is not. Every step reads either one row's numbers or a sum over all the rows,
and neither depends on which row is which. So without a positional encoding, the
decoder receives the same vector for I love LLMs and
LLMs love I in every block. It can only write the same translation for
both.
Step 4 is the one Section 3.4 called the point where the order is really lost. Up to the top of the encoder, every matrix has one row per English word, so the order at least survives as the row numbers. Step 4 adds those rows into one vector, and from then on there is no English row number to keep. The only thing that passes through a sum is the contents of the vectors being added. Those vectors come from \(\mathbf{C}\), and \(\mathbf{C}\) comes from \(\mathbf{X}\). So the position has to be inside \(\mathbf{X}\) before the encoder runs, which is Section 3.5. The other choice is to put it in where \(\mathbf{q}\cdot\mathbf{k}\) is computed, which is what RoPE does in Part 4.