The Transformer, Number by Number  ·  Part 2

Multi-Head Attention

One attention head gives every token a single budget to spend. That turns out to be the wrong shape for language, and the fix is to run several heads side by side on slices of the same vector.

Before you startThis page assumes Part 1: where \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) come from, how \(\mathbf{Q}\mathbf{K}^\top\) becomes an attention weight matrix, and what \(d_k\) counts. If you can read Part 1, Equation 1-1 and say what each symbol is, you are ready.

2.1Why one head is not enough

Everything in Part 1, Section 1.2 used a single attention weight matrix. It works. The paper does not use a single matrix, and the reason is worth spelling out, because that reason is the whole argument for the word multi-head in the title of this part.

Look again at the attention weight matrix in Part 1, Section 1.2.2 — the row for love, the middle one. That row is one budget, and it has to be spent on every relationship love takes part in at once. But a verb in a sentence is doing several unrelated jobs at the same time. It has a subject to its left. It has an object to its right. It agrees in tense with something, and it may be the answer to a question asked ten words earlier. Those are different relations, and there is no reason the same set of weights should be right for all of them.

Think for a moment Suppose the best weights for finding the subject are [0.8, 0.1, 0.1], and the best weights for finding the object are [0.1, 0.1, 0.8]. With exactly one row of weights available, what does the model end up learning?

Something in the middle — roughly [0.45, 0.1, 0.45] — which is the right answer to neither question. That is the real objection in the paper, and it is easy to state now that you have seen the numbers: a single head produces one weighted average per position, and averaging blurs together information that would have been more useful kept apart. The paper's own way of putting it is that several heads let the model attend to information from different representation subspaces at once. A representation subspace is just a slice of the vector — a subset of the columns — which the model is free to use for its own purpose. One head sees one slice and forms one opinion about the sentence. Eight heads see eight slices and form eight opinions.

✗ Common mistake Attention weights often are heaviest on the diagonal, and it makes a memorable story, so a popular claim spreads: a single head pays too much attention to itself, and multi-head exists to fix that. That is wrong. It is an observation about trained models, not the argument in the paper, and it is not always true. The paper's reason is the averaging one above. Treat the diagonal as an illustration, not as the justification.

So the design is: run attention several times in parallel, each time on its own slice of the vector, then join the results back together. Each parallel copy is called a head. Four symbols appear in the two equations below, so here they are first: \(h\) is the number of heads; the subscript \(i\) says which head, so \(\text{head}_i\) is head number \(i\); \(\mathbf{W}_i^Q\), \(\mathbf{W}_i^K\) and \(\mathbf{W}_i^V\) are the three projections that head \(i\) uses, and Section 2.2 shows where they come from; and \(\mathbf{W}^O\) is the matrix that recombines the heads at the end.

\[ \text{MultiHead}(\mathbf{Q},\mathbf{K},\mathbf{V})=\text{Concat}(\text{head}_1,\dots,\text{head}_h)\,\mathbf{W}^O,\tag{2-1} \] \[ \text{head}_i=\text{Attention}(\mathbf{Q}\mathbf{W}_i^Q,\,\mathbf{K}\mathbf{W}_i^K,\,\mathbf{V}\mathbf{W}_i^V) \tag{2-2} \]

Equation 2-2 takes three inputs, named \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) in the paper. In this part one matrix goes into all three:

\[ \mathbf{Q} = \mathbf{K} = \mathbf{V} = \mathbf{X} \]

where \(\mathbf{X}\) is the embedded sentence. So Equation 2-2 reads:

\[ \text{head}_i=\text{Attention}(\mathbf{X}\mathbf{W}_i^Q,\,\mathbf{X}\mathbf{W}_i^K,\,\mathbf{X}\mathbf{W}_i^V) \]

Part 1's \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) were not these three. Those were \(\mathbf{X}\mathbf{W}^Q\) and its two partners, built before the formula was applied. \(\mathbf{X}\mathbf{W}_1^Q\) is what Section 2.2.2 works out as \(\mathbf{Q}_1\).

Part 1 used one term without settling it, and the next equation needs it settled. \(d_{\text{model}}\) is the width of the vector the model carries for each token: one number per slot of that vector — a slot being where a number sits inside one vector, counting from 0. The width is the same at the input to every layer and at the output of every layer. The paper sets it to 512. \(d_{\text{model}}\) is the one size that does not change anywhere in the network, which is why the paper can name it once and then use it everywhere. Section 2.5 works out what forces it to stay fixed; for now, read it as "the width of the vector each token carries".

One warning about numbering: slots inside a vector count from 0, as above, and columns of a matrix count from 1, which is what everyone writing about matrices does. They point at the same places: column 1 of \(\mathbf{Q}\) holds slot 0 of every row. Part 3, Section 3.2 sets both counts out in full. From here on this part talks about columns, so every number you meet next to the word "column" starts at 1.

And the widths are tied to the head count, so that \(h\) heads cost about the same as one:

\[ d_q = d_k = d_v = \frac{d_{\text{model}}}{h} \tag{2-3} \]

With the paper's settings, \(d_{\text{model}} = 512\) and \(h = 8\), so each head works in 64 dimensions instead of 512.

Key takeaway

2.2What splitting actually splits

2.2.1What the split does

"Split into eight heads" sounds like the model builds eight of everything. It does not, and the order of operations is the part worth getting right, because it is also how the code is written. Figure 1 shows the split on the very numbers this section is about to compute; those numbers are worked out from scratch in Step 1 below, so read the figure now for the shape of the operation and come back to it once you have them.

Q 2 1 0 0 1 2 0 1 1 0 1 2 I love LLMs 3 × 4 columns 1–2 columns 3–4 Q1 Q2 2 1 1 2 1 0 I love LLMs 0 0 0 1 1 2 I love LLMs 3 × 2 3 × 2 K and V are cut exactly the same way. Nothing is copied and nothing is recomputed — the columns are simply dealt out to the heads. Row 1 of Q2 is all zeros — head 2 has no query for "I".
Figure 1. What splitting into heads means concretely, with \(d_{\text{model}} = 4\) and \(h = 2\). Each half sits directly under the columns it came from, and every number below already appears above.

\(\mathbf{X}\) is multiplied by \(\mathbf{W}^Q\), \(\mathbf{W}^K\) and \(\mathbf{W}^V\) to give full-width \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\). That step neither knows nor cares what \(h\) is — in real code you create a single \(\mathbf{W}^Q\) of shape \(d_{\text{model}} \times d_{\text{model}}\), not \(h\) separate ones. Only afterwards are the columns dealt out into \(h\) groups, and that split is the only place \(h\) enters at all.

✗ Common mistake The paper writes the heads apart, so it is tempting to read Equation 2-2 as the model stores \(h\) separate projection matrices, but that is wrong. It stores one. Slicing the columns of a single \(d_{\text{model}} \times d_{\text{model}}\) \(\mathbf{W}^Q\) into \(h\) groups gives exactly the \(\mathbf{W}_i^Q\) of Equation 2-2. The paper writes them apart so that it can name one head at a time. The code keeps them together, because one large matrix multiply is faster than \(h\) small ones.

Each head then runs Part 1, Equation 1-1 on its own slice, independently of the others. The \(h\) outputs are laid side by side again, and one final matrix \(\mathbf{W}^O\) mixes them back into a single vector per token (Figure 2).

X n × dmodel Q n × dmodel K n × dmodel V n × dmodel WQ WK WV head 1  ·  columns 1–2 Q1   K1   V1 n × dmodel/2 scaled dot-product attn head 2  ·  columns 3–4 Q2   K2   V2 n × dmodel/2 scaled dot-product attn Z1 n × dmodel/2 Z2 n × dmodel/2 Z concatenate n × dmodel output n × dmodel WO
Figure 2. Multi-head attention with \(h = 2\): one projection produces full-width Q, K and V, whose halves attend independently before being concatenated and passed through \(\mathbf{W}^O\). Each head reaches into all three matrices, which is why the lines cross — a head is a slice taken across Q, K and V together, not a slice of any one of them.

2.2.2The layer on three words

Words are one thing; digits are another. The rest of this section runs the whole layer on one sentence, starting from the embedded tokens, so you can follow every number here without going back to Part 1 — the formula each head runs is Part 1's, unchanged. The sentence is I love LLMs, and the model width is \(d_{\text{model}} = 4\) with \(h = 2\) heads.

The setup

First \(\mathbf{X}\), the embedded sentence. Three tokens, four numbers each, so 3 × 4. Row \(i\) is token \(i\):

\[ \mathbf{X}=\begin{pmatrix}0&0&1&1\\0&1&0&1\\1&0&0&1\end{pmatrix} \]

Then the three weight matrices that training produced, each 4 × 4. Note there is one of each, not one per head:

\[ \begin{aligned} \mathbf{W}^Q&=\begin{pmatrix}0&-1&1&1\\0&1&0&0\\1&0&0&-1\\1&1&0&1\end{pmatrix}\\[8pt] \mathbf{W}^K&=\begin{pmatrix}0&-1&0&1\\-1&1&0&0\\1&-1&1&-1\\1&1&0&1\end{pmatrix}\\[8pt] \mathbf{W}^V&=\begin{pmatrix}0&0&1&0\\1&0&0&1\\0&1&0&0\\1&0&0&0\end{pmatrix} \end{aligned} \]

Small whole numbers, chosen so the arithmetic stays in your head. A real model has decimals and 512 columns. Every step below would be identical.

These three are square only to keep the arithmetic small. In a real model \(\mathbf{W}^Q\) and \(\mathbf{W}^K\) are \(d_{\text{model}} \times d_k\) and \(\mathbf{W}^V\) is \(d_{\text{model}} \times d_v\); Part 1, Section 1.3.3 sets out which of those sizes are forced.

Step 1 — project to full width

Entry \((i,j)\) of \(\mathbf{X}\mathbf{W}^Q\) is row \(i\) of \(\mathbf{X}\) dotted with column \(j\) of \(\mathbf{W}^Q\). The top-left entry takes row 1 of \(\mathbf{X}\), which is \((0,0,1,1)\), against column 1 of \(\mathbf{W}^Q\), which is \((0,0,1,1)\):

\[ (0)(0)+(0)(0)+(1)(1)+(1)(1)=2 \]

Thirty-six such dot products fill in all three matrices:

\[ \begin{aligned} \mathbf{Q}=\mathbf{X}\mathbf{W}^Q&=\begin{pmatrix}2&1&0&0\\1&2&0&1\\1&0&1&2\end{pmatrix}\\[8pt] \mathbf{K}=\mathbf{X}\mathbf{W}^K&=\begin{pmatrix}2&0&1&0\\0&2&0&1\\1&0&0&2\end{pmatrix}\\[8pt] \mathbf{V}=\mathbf{X}\mathbf{W}^V&=\begin{pmatrix}1&1&0&0\\2&0&0&1\\1&0&1&0\end{pmatrix} \end{aligned} \]

All three are 3 × 4, full width. This step has not heard of \(h\) — it would produce exactly these matrices for one head or for eight.

Step 2 — deal the columns out

Now \(h\) arrives, and it arrives only here. With \(d_{\text{model}} = 4\) and \(h = 2\), each head gets two columns: head 1 takes columns 1 and 2, head 2 takes columns 3 and 4. Head 1's three slices:

\[ \begin{aligned} \mathbf{Q}_1&=\begin{pmatrix}2&1\\1&2\\1&0\end{pmatrix}\\[6pt] \mathbf{K}_1&=\begin{pmatrix}2&0\\0&2\\1&0\end{pmatrix}\\[6pt] \mathbf{V}_1&=\begin{pmatrix}1&1\\2&0\\1&0\end{pmatrix} \end{aligned} \]

And head 2's, taken from the same three matrices:

\[ \begin{aligned} \mathbf{Q}_2&=\begin{pmatrix}0&0\\0&1\\1&2\end{pmatrix}\\[6pt] \mathbf{K}_2&=\begin{pmatrix}1&0\\0&1\\0&2\end{pmatrix}\\[6pt] \mathbf{V}_2&=\begin{pmatrix}0&0\\0&1\\1&0\end{pmatrix} \end{aligned} \]

No arithmetic happened. Every number above already appears in \(\mathbf{Q}\), \(\mathbf{K}\) or \(\mathbf{V}\) — they have only been sorted into two piles. Each pile is 3 × 2, so \(d_k = d_v = 2\) inside a head, which is \(d_{\text{model}}/h\) exactly as Equation 2-3 promised.

Step 3 — head 1 attends on its slice

The same formula as always, on smaller matrices. Scores first:

\[ \mathbf{Q}_1\mathbf{K}_1^\top=\begin{pmatrix}4&2&2\\2&4&1\\2&0&1\end{pmatrix} \]

Now the divisor changes, and this is the one part of the computation the split really does alter. Inside a head \(d_k = 2\), not 4, so we divide by \(\sqrt{2} \approx 1.414\) rather than by 2:

\[ \frac{\mathbf{Q}_1\mathbf{K}_1^\top}{\sqrt{2}} =\begin{pmatrix}2.83&1.41&1.41\\1.41&2.83&0.71\\1.41&0&0.71\end{pmatrix} \]

Then softmax along each row. Take row 2, the row for love: \(e^{\sqrt{2}}=4.11\), \(e^{2\sqrt{2}}=16.92\) and \(e^{\sqrt{2}/2}=2.03\), which total 23.06 — the exponents are the unrounded scores, not the 1.41 printed above. Dividing each by that total, and doing the same for rows 1 and 3:

\[ \mathbf{A}_1=\begin{pmatrix}0.67&0.16&0.16\\0.18&0.73&0.09\\0.58&0.14&0.28\end{pmatrix} \]

Every number printed on this page is rounded to two decimals, which is why row 1 adds up to 0.99 rather than 1.

Step 4 — head 2 does the same on the other slice

Its scores:

\[ \mathbf{Q}_2\mathbf{K}_2^\top=\begin{pmatrix}0&0&0\\0&1&2\\1&2&4\end{pmatrix} \]

And after dividing by \(\sqrt{2}\) and applying softmax:

\[ \mathbf{A}_2=\begin{pmatrix}0.33&0.33&0.33\\0.14&0.28&0.58\\0.09&0.18&0.73\end{pmatrix} \]
Think for a moment Row 1 of \(\mathbf{A}_2\) is three equal weights. Look back at \(\mathbf{Q}_2\) and work out why before reading on.

Row 1 of \(\mathbf{Q}_2\) is \((0,0)\) — the query vector for I sits entirely in the half that went to head 1, so head 2 has nothing to ask with. Every score in that row is a dot product with a zero vector, so all three come out 0, and softmax over three equal numbers is a flat \(1/3\) each. Head 2 does not fail here; it abstains, spreading its attention evenly because it has no reason to prefer anything. Real models do this too, and heads that abstain on most inputs are exactly the ones that prune away cheaply.

Now compare the two weight matrices, which is the point of the whole exercise (Figure 3).

head 1  ·  columns 1–2 I love LLMs 0.67 0.16 0.16 0.18 0.73 0.09 0.58 0.14 0.28 I love LLMs 3 × 3 head 2  ·  columns 3–4 I love LLMs 0.33 0.33 0.33 0.14 0.28 0.58 0.09 0.18 0.73 I love LLMs 3 × 3 same sentence, same Q, K, V, different halves head 1 sends "LLMs" back to "I" (0.58) head 2 keeps "LLMs" on itself (0.73) neither row could hold both readings at once — that is the whole argument for h > 1
Figure 3. The two heads' attention weight matrices, from Step 3 and Step 4 above. Same sentence, same \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) — the only difference is which two columns each head was given.

Look at row 3, the row for LLMs. Head 1 spends 0.58 of it on I; head 2 spends 0.73 on LLMs itself. Those are two different readings of the same word in the same sentence, and a single head is one row of one matrix — it cannot hold both. Two heads can.

Step 5 — each head spends its weights on its own V slice

Every head finishes by multiplying its weights into its own 3 × 2 value slice:

\[ \begin{aligned} \mathbf{Z}_1=\mathbf{A}_1\mathbf{V}_1&=\begin{pmatrix}1.16&0.67\\1.73&0.18\\1.14&0.58\end{pmatrix}\\[8pt] \mathbf{Z}_2=\mathbf{A}_2\mathbf{V}_2&=\begin{pmatrix}0.33&0.33\\0.58&0.28\\0.73&0.18\end{pmatrix} \end{aligned} \]

Each is 3 × 2: three tokens, two numbers each, because a head only ever sees its own \(d_v = 2\) columns. Neither head has produced anything the next layer could use — both are half as wide as the model needs.

Step 6 — lay the heads side by side

Concatenation restores the full width. Nothing is added or averaged; the two slices are simply written next to each other:

\[ \operatorname{Concat}(\mathbf{Z}_1,\mathbf{Z}_2) =\begin{pmatrix}1.16&0.67&0.33&0.33\\1.73&0.18&0.58&0.28\\1.14&0.58&0.73&0.18\end{pmatrix} \]

Back to the shape 3 × 4, which is where \(\mathbf{X}\) started.

Step 7 — \(\mathbf{W}^O\) does the mixing

The final matrix is 4 × 4 here, in general \(hd_v \times d_{\text{model}}\):

\[ \mathbf{W}^O=\begin{pmatrix}0&0&1&1\\1&1&0&0\\0&1&-1&1\\1&0&0&0\end{pmatrix} \]

Take the first entry of the output: row 1 of the concatenation dotted with column 1 of \(\mathbf{W}^O\). That column is \((0,1,0,1)\), which picks the second number from head 1's slice and the second from head 2's:

\[ (1.16)(0)+(0.67)(1)+(0.33)(0)+(0.33)(1)=0.67+0.33=1.00 \]

One number, built from both heads at once. The matrix below prints it as 1.01, because it was computed from the unrounded numbers. Eleven more dot products finish the layer:

\[ \operatorname{MultiHead}(\mathbf{X},\mathbf{X},\mathbf{X})=\operatorname{Concat}(\mathbf{Z}_1,\mathbf{Z}_2) \mathbf{W}^O =\begin{pmatrix}1.01&1.01&0.83&1.50\\0.46&0.75&1.16&2.31\\0.75&1.31&0.41&1.87\end{pmatrix} \]

Three tokens in, three tokens out, four numbers each — the same 3 × 4 that \(\mathbf{X}\) had at the top of this section. That the shape survives the whole journey is not a coincidence, and Section 2.5 explains what forces it.

✗ Common mistake It is tempting to read this output as evidence that multi-head is better, but that is wrong. Nothing on this page shows it: the weight matrices were made up and the model was never trained. The claim multi-head attention actually makes is about what a trained model can represent — with one head, row 3 must choose between looking at I and looking at LLMs; with two, it does not have to. Whether that freedom pays off is a question for measurement, which is what Section 2.4 turns to.
Key takeaway

Try it (5 minutes): Compute row 2 of the concatenation yourself, from \(\mathbf{A}_1\), \(\mathbf{V}_1\), \(\mathbf{A}_2\) and \(\mathbf{V}_2\) above. You should get \((1.73,\ 0.18,\ 0.58,\ 0.28)\). Then say which two of those four numbers head 1 is responsible for.

✗ Common mistake Concatenation already produces the right width, so it is tempting to assume \(\mathbf{W}^O\) is bookkeeping, there to fix the shape, but that is wrong. The shape is already right after concatenation: \(h\) slices of width \(d_{\text{model}}/h\) come to \(d_{\text{model}}\). \(\mathbf{W}^O\) is there to mix. Without it, the output vector would be eight untouched slices sitting next to each other, and no later layer could combine what two heads found unless it did the mixing itself.

Try it (3 minutes): Take \(d_{\text{model}} = 512\) and \(h = 8\), with a sentence of \(n\) tokens. Write down the shape of \(\mathbf{Q}\) after the projection, the shape of one head's \(\mathbf{Q}_i\) after the split, and the shape after the concatenation. Then find where in that sequence the number 8 first gets involved — and where it disappears again.

2.3One wide head vs many narrow ones

A question that comes up constantly at this point: if the total width is the same either way, what does splitting actually buy you?

Start with what does not change. One head of width \(d_{\text{model}}\), two heads of width \(d_{\text{model}}/2\), three of width \(d_{\text{model}}/3\) — all three use the same number of parameters and emit an output of the same shape. From the outside they are interchangeable. For most of the computation they are also identical on the inside (Figure 4).

dmodel  —  identical in all three cases h = 1 head 1 dk = dmodel h = 2 head 1 head 2 dk = dmodel/2 h = 3 head 1 head 2 head 3 dk = dmodel/3 nothing differs until the split — h decides only how finely Q, K and V are cut
Figure 4. The same \(d_{\text{model}}\), cut three ways. Up to the moment of the split, one wide head and several narrow ones do exactly the same arithmetic.

That sets up a small test. Suppose I hand you \(\mathbf{Q}\), \(\mathbf{K}\) and \(\mathbf{V}\) and ask for the output. You cannot compute it — not because anything is missing from the diagram, but because I have not told you \(h\), and \(h\) decides where the cuts fall. At \(h = 2\) each head sees half the columns; at \(h = 3\), a third of them.

Why does that matter so much? Because each head builds its own attention weight matrix, independently, out of its own slice. One head gives you a single \(n \times n\) weighting of the sequence. Two heads give you two, and nothing forces them to agree (Figure 5).

h = 1 one pattern — every token gets exactly one weighting h = 2 a second pattern is free to look elsewhere h = 3 a third adds another view again more heads → more independent weightings — but each one gets a narrower dk, so bigger h is not always better
Figure 5. Three panels, one per head count: at \(h = 1\) the sequence gets a single weighting, at \(h = 2\) a second one is free to look elsewhere, at \(h = 3\) a third. More heads buy more independent weightings but give each one a narrower \(d_k\), so more heads and wider heads pull against each other, and Section 2.4 measures which way to lean.

That is the payoff, and it is the direct answer to the objection raised in Section 2.1. One head can go on tracking whatever that head finds most useful, while another learns to link a verb to the verb's object, or a noun to the determiner in front of the noun, or whatever else turns out to be worth a head of its own.

There is a second effect, and it is a cost rather than a benefit. Because \(h\) sets \(d_k\), \(h\) also caps how much a single head can express — and the cap is sharper than "less room for information." The cap is a hard ceiling on the set of patterns a head can produce at all, one that no amount of training lifts. Section 2.5.4 shows a head failing on a pattern you would expect to be easy, and Appendix A works out where the ceiling sits in general. The short version: a head's attention scores have rank at most \(d_k\), and a great many patterns need more rank than that. The rank of a matrix is the number of its rows that are genuinely different — rows that are not just other rows rescaled and added together. Appendix A, Rank, in one paragraph, gives the careful version.

✗ Common mistake \(h\) changes \(d_k\), and \(d_k\) sets the scaling factor \(\sqrt{d_k}\), so it is a natural chain of thought to conclude the softmax sees differently-scaled scores at different head counts. It reaches the wrong end. The division by \(\sqrt{d_k}\) is designed precisely to cancel that difference: under the assumption in Part 1, Section 1.3.2, the scaled scores have variance about 1 whatever \(d_k\) is. Scaling is what makes head counts comparable. The real cost of a large \(h\) is the rank limit above, not a scale mismatch.
Key takeaway

2.4How many heads

Two effects pull in opposite directions. More heads means more independent weightings; more heads also means a narrower \(d_k\) for each of them. Somewhere in between is a best answer, and the paper measured it rather than guessed (Figure 6).

The measurement is reported in BLEU, the standard score for machine translation: it compares what the model wrote against human translations of the same sentences, higher is better, and on this benchmark a difference of half a point is worth arguing about.

TABLE 3, ROW (A) + BASE  ·  computation held constant 25.8 25.4 25.0 BLEU on the English-to-German development set 24.9 25.5 25.8 25.8 25.4 h = 1 h = 4 h = 8 h = 16 h = 32 dk 512 dk 128 dk 64 dk 32 dk 16 −0.9 BLEU −0.4 a tie, not a peak
Figure 6. Table 3 of the paper — row (A), plus the base row that supplies the \(h = 8\) point: BLEU on the English-to-German development set as the head count varies, with total computation held constant. The drop is lopsided — one head costs 0.9 BLEU, thirty-two costs 0.4.

The measured values are 24.9 at \(h = 1\), 25.5 at \(h = 4\), 25.8 at both \(h = 8\) and \(h = 16\), and 25.4 at \(h = 32\). Three things are worth reading off that list.

The first is that single-head attention is 0.9 BLEU worse than the best setting, which is a large gap for this benchmark. That gap is the strongest evidence in the paper for the argument in Section 2.1.

The second is that the two ends are not symmetric. Too many heads also hurts, but by 0.4 rather than 0.9 — less than half as much. Going too narrow is a milder failure than going too wide.

The third is that \(h = 8\) and \(h = 16\) score the same. Eight is not a peak the authors found. Eight is a point on a flat stretch, and the authors took the cheaper end of that stretch. If you have seen eight described as an optimum, this table is where the claim comes from, and it does not quite say that.

Key takeaway

2.5The Math Behind It

This section is optional on a first pass. It answers two questions the earlier sections raised without settling: why did 3 × 4 go into the layer and 3 × 4 come out, and why does a narrow head have less to say than a wide one?

2.5.1What \(d_{\text{model}}\) actually is

Section 2.1 gave a working definition: the width of the vector the model carries for each token. That is correct, but it undersells how rigid the number is. A more useful picture is a stream — one vector per token, flowing from the embedding layer to the output, with the same width at every point along the way (Figure 7).

the residual stream  ·  width locked at dmodel embeddings dmodel x dmodel x′ dmodel x″ dmodel multi-head attention inside: dk = dv = 64 free to be anything out back in feed-forward inside: dff = 2048 free to be anything out back in every box on this line is dmodel wide a sublayer may be as wide as it likes inside — it must hand back the width it was given
Figure 7. The residual stream. Every box on the vertical line carries the same width; the sublayers hanging off it may use any width they like internally, but each must return what it was handed.

Sublayers do not sit in that stream so much as hang off it. Each one reads the current vector, computes something, and adds its result back. In the paper's notation, with \(x\) the vector arriving and \(\text{Sublayer}\) whatever that sublayer computes:

\[ \text{output} = \text{LayerNorm}\bigl(x + \text{Sublayer}(x)\bigr)\tag{2-4} \]

Two names in that formula belong to later parts, and this is the only place you will meet them before then. \(\text{LayerNorm}\) puts a row on one scale without changing how many numbers it holds; Part 5, Section 5.3 builds it. The feed-forward sublayer is the other thing that hangs off the stream, and Part 5, Section 5.4 builds that. For this section all you need from either is that neither changes the width.

✗ Common mistake It is tempting to read \(d_{\text{model}}\) as the embedding size. That part is true: \(d_{\text{model}}\) is the embedding size. What the reading misses is that being the embedding size is a consequence, not the definition. The embedding layer has to emit \(d_{\text{model}}\) numbers because it is feeding the stream, not the other way round. Getting it backwards makes the next subsection look like an arbitrary rule instead of a forced one.

2.5.2Why the width cannot change

The plus sign in Equation 2-4 is the whole answer. Addition of vectors is entrywise: the first number pairs with the first, the second with the second, and so on to the end. It is only defined when both sides have the same number of entries.

Take \(d_{\text{model}} = 4\) and one token, so \(x\) is 1 × 4. Suppose the sublayer returns 1 × 4 as well. Then the addition goes through, slot by slot:

\[ \begin{array}{r|cccc} x & 0 & 1 & 0 & 1\\ \text{Sublayer}(x) & 0.46 & 0.75 & 1.16 & 2.31\\\hline \text{sum} & 0.46 & 1.75 & 1.16 & 3.31 \end{array} \]

Now suppose the sublayer had been built to return three numbers instead of four — a perfectly reasonable thing for a matrix multiply to do, and nothing inside attention prevents it. The same table:

\[ \begin{array}{r|cccc} x & 0 & 1 & 0 & 1\\ \text{Sublayer}(x) & 0.46 & 0.75 & 1.16 & -\\\hline \text{sum} & 0.46 & 1.75 & 1.16 & \;? \end{array} \]

The first three columns are fine. The fourth has a number on top and nothing underneath. There is no sensible value to write: dropping it would silently discard a quarter of what the residual path was carrying, and padding with zero would be an invented number pretending to be a computed one. The operation is simply not defined, and in code it raises a shape error rather than doing something clever.

Think for a moment Attention on its own never runs into this. In Section 2.2 the layer took 3 × 4 to 3 × 4, but look at \(\mathbf{W}^O\) — it was 4 × 4. What in the mathematics of attention forced that second 4 to be a 4?

Nothing did. \(\mathbf{W}^O\) could have been 4 × 3 or 4 × 7, and every matrix multiplication in Section 2.2 would still have had matching shapes. Attention is perfectly happy to change the width. The constraint comes from outside it — from the addition in Equation 2-4, which the attention sublayer has to feed.

Three other parts of the architecture are fixed by this same requirement:

Key takeaway

2.5.3What is free to vary

The width is locked between sublayers. Inside one, nothing is. This is worth stating plainly, because it is easy to over-generalise from the previous subsection and assume every dimension in the model must equal \(d_{\text{model}}\).

So a sublayer may be as wide or as narrow as it likes internally. It just has to hand back the width it was given.

✗ Common mistake It is tempting to conclude that a pretrained 300-dimensional word embedding forces \(d_{\text{model}} = 300\), but that is wrong. Put a learned 300 × 512 matrix immediately after the lookup and the stream runs at 512 from there on. Modern large models do exactly this in reverse as well, deliberately decoupling the embedding width from the stream width to avoid spending an enormous number of parameters on a large vocabulary — the fixed list of words the model can read, named in Part 6, Section 6.4.1.

One honest caveat. What must stay constant is the width of the residual stream, and \(d_{\text{model}}\) is simply the name the paper gives that width. An architecture built without residual connections could change width from layer to layer, and some do. Within the Transformer, and within essentially everything built on it since, the two are the same number because the architecture is built that way.

Try it (3 minutes): You are building a model with \(d_{\text{model}} = 768\) and want \(h = 12\) heads. Write down \(d_k\), the shape of one head's \(\mathbf{W}_i^Q\), and the shape of \(\mathbf{W}^O\). Then say which of those three numbers you could change without touching anything outside the attention sublayer.

2.5.4Why a narrow head can express less

Section 2.3 claimed that a small \(d_k\) limits what a head can represent, and left the reason here. This subsection makes that precise.

Here is what we want to test: with enough training, a head can produce any attention pattern you want. Is that true?

To show it is not, we only need to find one pattern the head cannot produce. One is enough — if the head fails even once, "any pattern" is wrong. So we build the smallest head there is and look for something it cannot do.

One note on letters before we start, because two of them are about to sit next to each other. \(\mathbf{Q}_i\) and \(\mathbf{K}_i\) are one head's slices, and the subscript \(i\) names the head, exactly as in Equation 2-2. Rows inside those slices are numbered with \(r\), and columns with \(j\). The rule is worth saying out loud: a capital letter with a subscript names a head, and a lower-case letter with a subscript names one row or one column inside it. So \(q_r\) is one token's query and \(k_j\) is one token's key.

The smallest possible head

Set \(d_k = 1\). The head's slices keep the shape they had in Section 2.2, only one column wide instead of two, so \(\mathbf{Q}_i\) and \(\mathbf{K}_i\) are both 3 × 1. Suppose training produced these keys, with the rows in the usual order I, love, LLMs:

\[ \mathbf{K}_i=\begin{pmatrix}2\\0\\1\end{pmatrix} \]

Leave the queries as unknowns, since what we want to know is what the head could learn:

\[ \mathbf{Q}_i=\begin{pmatrix}q_1\\q_2\\q_3\end{pmatrix} \]

Now multiply. A 3 × 1 against a 1 × 3 gives the full 3 × 3 score matrix:

\[ \mathbf{Q}_i\mathbf{K}_i^\top=\begin{pmatrix}2q_1&0&q_1\\2q_2&0&q_2\\2q_3&0&q_3\end{pmatrix} \]

Read the rows. Every one is \(q_r\) times the same triple \((2, 0, 1)\), which is \(\mathbf{K}_i^\top\). Nine entries, but only three free numbers, and the shape of every row fixed in advance. The head's entire freedom for row \(r\) is the single value \(q_r\). Sweep it and watch the row change (Figure 8):

dk = 1  ·  keys Ki = (2, 0, 1)T  ·  every row is q × KiT, then softmax I love LLMs 0.87 0.02 0.12 0.67 0.09 0.24 0.33 0.33 0.33 0.09 0.67 0.24 0.02 0.87 0.12 q = 2 q = 1 q = 0 q = −1 q = −2 one knob, one sweep The boxed column never wins. Its weight peaks at 0.33, in the three-way tie at q = 0, and falls away in both directions. No query makes this head attend mostly to "LLMs" — not a training failure, an impossibility. Only argmax(k) and argmin(k) can ever take the top weight.
Figure 8. All five rows come from the same three keys; only \(q\) changes. With \(d_k = 1\) the reachable rows form a one-parameter family, and one of the three tokens is locked out of ever receiving the most attention.

Large positive \(q\) concentrates on I; large negative \(q\) concentrates on love; \(q = 0\) gives a flat row. But LLMs never takes the top weight. Its best showing is 0.33, in the three-way tie, and it drops away in both directions.

A pattern it cannot produce

Now pick a target that needs exactly what the sweep cannot deliver. Take the plainest pattern imaginable — every token attending mostly to itself:

\[ \mathbf{A}=\begin{pmatrix}0.8&0.1&0.1\\0.1&0.8&0.1\\0.1&0.1&0.8\end{pmatrix} \]

Row 3 of this target puts its largest weight on LLMs. The sweep says no value of \(q_3\) does that. So this head cannot produce this \(\mathbf{A}\) — and "with enough training, a head can produce any attention pattern you want" is wrong.

But look closely at what we just showed. It was one head, with keys I chose, failing on one target — which does not yet rule out a cleverer head. Appendix A closes that gap. It proves that no choice of keys helps, that the same limit applies at every \(d_k\), and that producing this \(\mathbf{A}\) needs a score matrix of rank at least \(n - 1\). Nothing later in this part or in Part 3 depends on the appendix, so you can safely skip it on a first reading.

Try it (3 minutes): Take \(d_k = 1\) again with \(\mathbf{K}_i\) as above, and find the \(q\) that gives love a weight above 0.95. Then try to find one that gives LLMs a weight above 0.34, and convince yourself why the search fails.

AAppendix: how far the rank bound reaches

This was pulled out of Section 2.5.4 so that a first reading runs without stopping. Nothing later depends on it. It answers the three questions the worked example left open — whether the keys mattered, whether \(d_k = 1\) mattered, and how much rank the target \(\mathbf{A}\) really demands — and ends with the whole argument written out as one list.

Question 1 — was it the keys I picked?

The keys \((2, 0, 1)^\top\) were chosen by hand. Maybe a better-trained head picks keys that work.

It cannot. Row \(r\) of \(\mathbf{Q}_i\mathbf{K}_i^\top\) is \(q_r\bigl(\mathbf{K}_i\bigr)^{\!\top}\) — a single scalar multiplying one fixed row. A positive scalar preserves the order of that row's entries and a negative one reverses it, so no scalar can promote an entry that is neither the largest nor the smallest. Writing \(k_j\) for the \(j\)-th entry of \(\mathbf{K}_i\), the winner of row \(r\) is therefore the largest key when \(q_r > 0\), the smallest when \(q_r < 0\), and nobody at all when \(q_r\) is exactly 0:

\[ \text{top-attended token of row } r= \begin{cases} \arg\max_j k_j, & q_r > 0\\[4pt] \arg\min_j k_j, & q_r < 0\\[4pt] \text{no winner — the row is flat}, & q_r = 0 \end{cases} \]

Two candidates, ever — plus a flat row when \(q_r\) is exactly 0. With three tokens that leaves one permanently locked out — whichever key lands in the middle. With 512 tokens it leaves 510 locked out. The specific keys make no difference; only their order matters, and every order has exactly one largest and one smallest.

Question 2 — does this only happen at \(d_k = 1\)?

No. And the way to see it is to redo the same calculation one size up.

Set \(d_k = 2\), so \(\mathbf{Q}_i\) and \(\mathbf{K}_i\) are both 3 × 2. Write the columns of \(\mathbf{Q}_i\) as \(\mathbf{q}^{(1)}, \mathbf{q}^{(2)}\) and the columns of \(\mathbf{K}_i\) as \(\mathbf{k}^{(1)}, \mathbf{k}^{(2)}\). Multiplying out:

\[ \mathbf{Q}_i\mathbf{K}_i^\top\;=\;\mathbf{q}^{(1)}\bigl(\mathbf{k}^{(1)}\bigr)^{\!\top}\;+\;\mathbf{q}^{(2)}\bigl(\mathbf{k}^{(2)}\bigr)^{\!\top} \]

Each term is the same object we met at \(d_k = 1\): a column times a row, every row of it a scalar multiple of one fixed row. So at \(d_k = 2\), \(\mathbf{Q}_i\mathbf{K}_i^\top\) is a sum of two such matrices, and row \(r\) is

\[ \bigl(\mathbf{Q}_i\mathbf{K}_i^\top\bigr)_{r,:}\;=\;q_{r1}\bigl(\mathbf{k}^{(1)}\bigr)^{\!\top}\;+\;q_{r2}\bigl(\mathbf{k}^{(2)}\bigr)^{\!\top} \]

At \(d_k = 1\) row \(r\) was \(q_r\bigl(\mathbf{K}_i\bigr)^{\!\top}\) — a scalar multiple of a single row. At \(d_k = 2\) it is a linear combination of two rows, with \(q_{r1}\) and \(q_{r2}\) as the coefficients — a linear combination of two rows means: multiply each row by a number, then add the two results. The pattern continues: at a general \(d_k\), row \(r\) is a linear combination of \(d_k\) fixed rows, namely the rows of \(\mathbf{K}_i^\top\).

So the quantity that governs what a head can express is the number of rows in that combination. That number has a name: the rank of \(\mathbf{Q}_i\mathbf{K}_i^\top\).

Rank, in one paragraph

The rank of a matrix is the number of its rows that are linearly independent — independent meaning no one of them equals a linear combination of the others. Three rows that are all scalar multiples of a single row give rank 1. Three rows that are linearly independent give rank 3. Rank counts independent rows and nothing else: multiplying every entry of a matrix by 1000 leaves its rank unchanged. If this is new, 3Blue1Brown's Essence of Linear Algebra builds the picture visually and MIT's 18.06 covers it in full.

With that term available, the calculation above generalises in one line. For any \(\mathbf{X}\) with \(d\) columns and any \(\mathbf{Y}\) with \(d\) rows, row \(r\) of \(\mathbf{X}\mathbf{Y}\) is

\[ (\mathbf{X}\mathbf{Y})_{r,:}\;=\;x_{r1}\,\mathbf{Y}_{1,:}\;+\;x_{r2}\,\mathbf{Y}_{2,:}\;+\;\dots\;+\;x_{rd}\,\mathbf{Y}_{d,:} \]

where \(\mathbf{Y}_{1,:}\) through \(\mathbf{Y}_{d,:}\) are the rows of \(\mathbf{Y}\), and \(x_{r1}\) through \(x_{rd}\) are the entries of row \(r\) of \(\mathbf{X}\). Every row of \(\mathbf{X}\mathbf{Y}\) is therefore a linear combination of the same \(d\) rows, so \(\mathbf{X}\mathbf{Y}\) has at most \(d\) linearly independent rows:

\[ \operatorname{rank}(\mathbf{X}\mathbf{Y})\;\le\;d\tag{2-5} \]

Apply that with \(\mathbf{X} = \mathbf{Q}_i\) and \(\mathbf{Y} = \mathbf{K}_i^\top\), where \(d = d_k\):

\[ \operatorname{rank}\bigl(\mathbf{Q}_i\mathbf{K}_i^\top\bigr)\;\le\;d_k\tag{2-6} \]

One at \(d_k = 1\), two at \(d_k = 2\), and so on. That is the ceiling. Now for the floor. Write \(\mathbf{S}\) for a score matrix — whatever a head hands to softmax. How much rank must \(\mathbf{S}\) have, if softmax is to turn it into \(\mathbf{A}\)?

Question 3 — how much rank does a score matrix need in order to produce \(\mathbf{A}\)?

Work backwards. Rather than asking what this head can build, ask what anything would have to build in order to produce \(\mathbf{A}\), the diagonal target above with 0.8 on each diagonal entry and 0.1 elsewhere.

Softmax turns a score row into a weight row by computing \(e^{s_{rj}}\) and dividing by the total of those exponentials. Set that equal to \(A_{rj}\) and take the log of both sides:

\[ s_{rj}=\log A_{rj}+c_r\tag{2-7} \]

where \(c_r=\log\sum_j e^{s_{rj}}\) — the log of the total the softmax divided by, one number for the whole of row \(r\). Read it as: the required scores are \(\log \mathbf{A}\), plus the same constant added along each row. Nothing else works, and nothing else is needed.

So how much rank does \(\log \mathbf{A}\) have? All three. Writing \(a = \log 0.8 - \log 0.1 = 2.079\) and \(b = \log 0.1 = -2.303\),

\[ \log \mathbf{A} \;=\; a\mathbf{I} + b\mathbf{J} \]

with \(\mathbf{I}\) the identity and \(\mathbf{J}\) the all-ones matrix. Then \(\det(\log \mathbf{A}) = a^2(a + 3b) = -20.9\), and a non-zero determinant means full rank, so \(\operatorname{rank}(\log \mathbf{A}) = 3\).

And the constants? The added term \(\mathbf{c}\mathbf{1}^\top\) has every column equal to \(\mathbf{c}\), so all its columns are scalar multiples of one another and \(\operatorname{rank}(\mathbf{c}\mathbf{1}^\top) = 1\). What we need next is that adding a rank-1 matrix cannot drop the rank by more than 1. That takes two short steps.

Step A: ranks are subadditive. Row \(r\) of \(\mathbf{X} + \mathbf{Y}\) is row \(r\) of \(\mathbf{X}\) plus row \(r\) of \(\mathbf{Y}\). Every row of \(\mathbf{X}\) is a linear combination of some \(\operatorname{rank}(\mathbf{X})\) independent rows, and likewise for \(\mathbf{Y}\), so every row of \(\mathbf{X} + \mathbf{Y}\) is a linear combination of those two sets pooled together. Pooling them gives at most \(\operatorname{rank}(\mathbf{X}) + \operatorname{rank}(\mathbf{Y})\) rows to draw on, so

\[ \operatorname{rank}(\mathbf{X}+\mathbf{Y})\;\le\;\operatorname{rank}(\mathbf{X})+\operatorname{rank}(\mathbf{Y})\tag{2-8} \]

Step B: run it backwards. We want a lower bound on \(\operatorname{rank}(\mathbf{S})\), and Step A only gives upper bounds — so apply it to a sum that produces \(\log \mathbf{A}\) rather than \(\mathbf{S}\). Since \(\mathbf{S} = \log \mathbf{A} + \mathbf{c}\mathbf{1}^\top\), rearranging gives \(\log \mathbf{A} = \mathbf{S} + \bigl(-\mathbf{c}\mathbf{1}^\top\bigr)\), and negating a matrix leaves its rank alone. Step A applied to that sum reads

\[ \underbrace{\operatorname{rank}(\log \mathbf{A})}_{=\,3}\;\le\;\operatorname{rank}(\mathbf{S})+\underbrace{\operatorname{rank}(\mathbf{c}\mathbf{1}^\top)}_{=\,1} \]

which rearranges to \(\operatorname{rank}(\mathbf{S}) \ge 3 - 1 = 2\). That is the floor, and it holds for every \(\mathbf{S}\) that produces \(\mathbf{A}\).

Think for a moment Step B works because subtracting \(\mathbf{c}\mathbf{1}^\top\) is as legitimate a move as adding it. Where would the argument break if the perturbation had rank 3 instead of rank 1?

The bound would become \(\operatorname{rank}(\mathbf{S}) \ge 3 - 3 = 0\), which says nothing at all. A low-rank perturbation is what makes this useful: the smaller the thing you add, the less it can disturb the rank of what you started with, and a per-row constant is the smallest non-trivial perturbation there is.

Ceiling 1, floor 2. They do not meet — the same conclusion Figure 8 reached by hand, now with a reason that does not depend on which keys were learned. Figure 9 puts the two numbers on one axis.

rank of the score matrix  ·  three-token sentence 1 2 3 floor: every score matrix that produces A needs rank 2 dk = 1 ceiling 1 dk = 2 ceiling 2 cannot reach clears the floor clearing the floor is necessary, not sufficient — it says the target is not ruled out, while falling below it rules the target out for good
Figure 9. Ceiling against floor: a head reaches every rank up to \(d_k\), while every score matrix that produces \(\mathbf{A}\) sits at rank 2 or above. At \(d_k = 1\) the two do not meet, which is what rules \(\mathbf{A}\) out.

The same reasoning works for a sentence of any length. When \(\log \mathbf{A}\) has full rank, every score matrix that produces \(\mathbf{A}\) has rank at least \(n - 1\). So a head whose \(d_k\) is smaller than that cannot produce \(\mathbf{A}\) at all.

The whole argument in one place

That took a while, so here is the path we walked, with the notation stripped out.

  1. We asked whether a head can produce any attention pattern, given enough training.
  2. We built the smallest head possible, \(d_k = 1\), and left its queries as unknowns so that "any training outcome" was covered. Its score matrix turned out to be \(q_r\) times one fixed row of keys — nine numbers, but only three of them free (2.5.4, The smallest possible head).
  3. Sweeping \(q_r\) showed that only two tokens can ever win a row: the one with the largest key and the one with the smallest. Any token whose key falls in between never wins (Figure 8).
  4. We picked a target \(\mathbf{A}\) — each token attending mostly to itself — whose third row needs precisely one of those in-between tokens to win. The head cannot do it, so "any pattern" is already false (2.5.4, A pattern it cannot produce).
  5. Were the keys the problem? No. A scalar multiplying a fixed row either preserves the order of its entries or reverses it, and every set of keys has exactly one largest entry and one smallest (Question 1).
  6. Was \(d_k = 1\) the problem? Redoing the calculation at \(d_k = 2\) showed row \(r\) is a linear combination of two fixed rows instead of a multiple of one, and in general of \(d_k\) rows. That count is the rank, and \(\operatorname{rank}(\mathbf{Q}_i\mathbf{K}_i^\top) \le d_k\). That is a ceiling (Equation 2-5, Equation 2-6).
  7. How high is the floor? Running softmax backwards, any \(\mathbf{S}\) with \(\operatorname{softmax}(\mathbf{S}) = \mathbf{A}\) equals \(\log \mathbf{A} + \mathbf{c}\mathbf{1}^\top\) (Equation 2-7). Since \(\operatorname{rank}(\log \mathbf{A}) = 3\) and \(\operatorname{rank}(\mathbf{c}\mathbf{1}^\top) = 1\), subadditivity (Equation 2-8) gives \(\operatorname{rank}(\mathbf{S}) \ge 2\).
  8. Ceiling 1, floor 2. No head with \(d_k = 1\) produces \(\mathbf{A}\), whatever it learned — and for a sentence of length \(n\) the same three steps give a floor of \(n - 1\) (Figure 9).

The shape worth remembering is step 7. Everywhere else we reasoned forwards, about what the head can do. There we reasoned backwards, about what \(\mathbf{A}\) demands. The impossibility only becomes visible when the two are put side by side.

✗ Common mistake It is tempting to turn Equation 2-6 into \(\mathbf{A}\) has rank at most \(d_k\), but that is wrong. Those are two different matrices. Equation 2-6 bounds \(\mathbf{S}\), the scores. \(\mathbf{A}\) is what comes out of softmax afterwards, and softmax is not linear, so it can raise the rank on the way through. \(\mathbf{A}_1\) back in Section 2.2 came from a head with \(d_k = 2\), yet its three rows are independent, so \(\operatorname{rank}(\mathbf{A}_1) = 3\) — higher than the bound on the scores that produced it.

Equation 2-6 still decides which \(\mathbf{A}\) a head can reach, just by a longer route. A head gets to \(\mathbf{A}\) in two moves: it builds \(\mathbf{S}\), then softmax turns \(\mathbf{S}\) into \(\mathbf{A}\). Question 3 showed that every \(\mathbf{S}\) that would produce our diagonal \(\mathbf{A}\) has rank 2 or more. Equation 2-6 says a head with \(d_k = 1\) can only build \(\mathbf{S}\) of rank 1. The first move already fails, so the second never happens — and none of that required knowing the rank of \(\mathbf{A}\) itself.

Now the numbers that matter in practice. At \(d_{\text{model}} = 768\) and \(h = 12\), each head gets \(d_k = 64\). On a 512-token context that is a 512 × 512 score matrix confined to rank 64, against an arbitrary target that would want something near 511. Push to \(h = 64\) and \(d_k\) falls to 12. The heads have not become worse at their job; the job has become impossible to do in full.

It is worth being clear about what rank is measuring here, because the intuition it rewards is not the obvious one. Rank measures how little the rows of a matrix repeat one another: the more the rows say the same thing, the lower the rank. A target where every token attends to the same position has identical rows and rank 1 — crude to look at, but cheap to build. Our diagonal target \(\mathbf{A}\), where every token attends to itself, looks far more ordinary but is the most expensive pattern a 3 × 3 matrix can hold: no row repeats another.

So the bound is severe in principle. An arbitrary target for a 512-token sentence wants a score matrix of rank near 511, and \(d_k = 64\) supplies 64. Whether that gap matters in practice is a separate question, and the answer is less alarming than the arithmetic suggests.

✗ Common mistake \(d_k = 64\) is far below any realistic \(n\), so it is tempting to conclude that real models must therefore be broken. That is wrong. The attention patterns language actually needs are highly structured — attend to the previous token, to the matching bracket, to the subject of the verb — and structured patterns are close to low rank already. The bound says arbitrary patterns are out of reach, not useful ones. It becomes a real constraint when \(h\) is pushed high enough that \(d_k\) falls below what even structured patterns need, which is one reason some later architectures set \(d_k\) independently instead of tying it to \(d_{\text{model}}/h\).