The Transformer, Number by Number  ·  Part 0

How a Sentence Becomes Numbers

Every calculation in Parts 1 to 8 starts from the same matrix: the sentence I love LLMs, already turned into three rows of four numbers each. This part is where those rows come from. It starts from eleven characters and finishes with the matrix, and every step in between is a count you can do with a pencil.

By the end you'll be able to cut a sentence into tokens by hand. You'll be able to say why LLMs comes out as two pieces rather than one. You'll be able to run the loop that builds a vocabulary, counting every pair yourself. And you'll be able to say what a token id is, and why the model in the paper this series works through — Attention Is All You Need, Vaswani and others, 2017 — knows about 37,000 tokens rather than every word in two languages.

You can read this before Part 1 or after Part 8. It needs no weight matrices and no calculus, and nothing in Parts 1 to 8 depends on having read it. A few sentences borrow a fact from a later part, and each one hands you the fact on the spot, so you are never sent away to follow the argument. Section 0.1 borrows how wide a row is and what the model's twelve blocks weigh. Section 0.2 borrows one fact about how self-attention works. Section 0.3 borrows the size of the paper's shared English-to-German vocabulary. Section 0.5 borrows the size of the encoder stack, what the decoder needs its three markers for, and the fact that the numbers in \(\mathbf{E}\) are learned rather than chosen. Section 0.6 borrows what the model produces each time it writes.

The example this series runs on. The transformer translates I love LLMs into J’aime les grands modèles. Both sentences appear here as text, before anything has been multiplied by anything. Figure 1 is the whole part in one picture, and every stage in it is built by the section named beside it.

FROM TEXT TO THE ROWS OF X the four stages of this part, and the section that builds each one I love LLMs the text a person types cut at spaces, then apply the merge list · Sections 0.3 and 0.4 I love LLM s four tokens one word, two tokens number the vocabulary · Section 0.5 0 28 21 14 four token ids take one row out of E for each id · Section 0.5.2 four rows: the matrix X 4 × d_model · the matrix this part builds Parts 1 to 8 simplify to one row per word, so they start from three · Section 0.5.2
Figure 1. The four stages of this part: the sentence as text, the four tokens it is cut into, the four ids of those tokens, and the four rows they pull out of E. The word LLMs arrives as two tokens, which is what the rest of this part explains.

0.1A row for every word, and none for the word you have not met

Open a paper dictionary and every word has its own entry. The obvious plan for a model is the same one. Give every word a row of numbers, then turn a sentence into the rows of its words. I love LLMs becomes three rows of four numbers, 3 × 4, and that is the matrix every later part starts from.

A large wooden rubber stamp printing the word INTELLIGENCE from an overflowing box of word stamps
A whole-word rubber stamp: one press prints a whole word like intelligence, but a fixed box cannot hold a stamp for every word anyone might write.

Think of a box of rubber stamps. Each stamp prints one piece of text, and you write a sentence by pressing stamps in order. How many times you press is how long the sentence is, and each press lands on one position. Now the trade-off, and it runs through this whole part. The box holds a fixed number of stamps, chosen in advance. Fill it with big stamps and each press prints a lot, so sentences come out short — but a fixed number of big stamps cannot cover everything anybody might write. Fill it with small stamps and nothing is ever unprintable, because you can spell anything out — but every sentence takes far more presses. You cannot have both. Figure 2 puts the two ends side by side.

ONE TRADE-OFF, TWO ENDS the box holds a fixed number of stamps; how big you make them decides everything else BIG STAMPSone stamp per whole word the boxlovegrandmodel… the presses I love LLMs 3 presses a fixed number of big stampscannot cover everything anybody writes SMALL STAMPSone stamp per character the boxIlovems… the presses I ␣ l o v e ␣ L L M s 11 presses nothing is ever unprintable, butevery sentence takes far more presses Section 0.3 puts the answer between these two ends
Figure 2. The same sentence under the two extremes: one stamp per whole word takes three presses but needs a box no fixed size can fill, and one stamp per character can print anything but takes eleven presses.
position — where a word sits in the sentence, counting from 0. Part 3, Section 3.2 defines it against the rows of \(\mathbf{X}\).

Take the big end first — one stamp per whole word — and count what it costs.

Think for a moment A hundred thousand words is a modest English vocabulary — a desk dictionary holds more. The paper's base model has \(d_{\text{model}} = 512\), so each row is 512 numbers. How many numbers is the whole table?

The table holds \(100{,}000 \times 512 = 51{,}200{,}000\) numbers. The twelve blocks that do the work come to about 44 million between them. The lookup table alone would be the biggest thing in the model, and it would still leave sentences it cannot read.

The sentences it cannot read are the ones with a word you have not met. Your surname is probably not in any hundred-thousand-word list. Nor is a product name from last month, nor a word borrowed from another language, nor a typing mistake. Each of those arrives at a model with one row per word and finds nothing. The usual repair is a single row meaning unknown, shared by every word that is not in the table. That gives all of those words the same row of numbers. A translator that gives your surname and a typing mistake the same row has stopped translating and started guessing.

There is a third cost, and it is the one that decides the design. English writes walk, walks, walked and walking. One stamp per word buys four rows there, and the model has to learn separately, four times over, that all four are about walking. French is worse: grand, grande, grands and grandes are four rows for one adjective, and our own sentence uses one of them. Figure 3 puts the three costs side by side.

1 · TOO BIGthe lookup table51,200,000 numbersagainst the twelve blocks that do the work:44,101,632 numbers, all twelve together2 · A WORD IT HAS NEVER METKowalczyk→no rowthe usual repair is one shared rowmeaningunknownso a surname, a typo and a foreignword all arrive as the same vectora translator that does thatis not translating3 · FOUR ROWS FOR ONE IDEAwalkits own rowwalksits own rowwalkedits own rowwalkingits own rowthe model has to learn four timesover that all four are about walkingFrench is worse: grand, grande,grands, grandes — four for one adjective
Figure 3. The three costs of one row per word, side by side: a table larger than the twelve blocks that do the work, a word the table has never met, and four rows spent on one idea. Byte-pair encoding is the answer to all three at once.

So what is on the list does not have to be a whole word. Whatever is on it is called a token, and the list itself is the vocabulary — the same name Part 6, Section 6.4.1 already gave it. A token is often a whole word, but it does not have to be.

Key takeaway

Try it (2 minutes): Take the last message you sent to anyone. Count the words that a fixed English list of 100,000 would not hold — names, abbreviations, typing mistakes, words from another language. In most real messages the count is not zero.

0.2A row for every character, and a sentence four times longer

Turn the plan over. Give every character a row instead of every word. The box of stamps now holds letters: one for a, one for b, one for the space. English text needs about a hundred of them once capitals, digits and punctuation are counted. Our two sentences use only eighteen different characters between them.

A compartmentalised tray of miniature character stamps being pressed letter by letter
Single-character stamps: a small box easily holds every letter, but spelling out a sentence takes a separate press for every character.

That plan can never meet an unknown word. Your surname is made of letters the box already holds, and so is a typing mistake, and so is LLMs. Nothing gets left out and nothing collapses into a shared unknown row. The unknown-word problem from Section 0.1 is gone completely.

So count what it costs. I love LLMs is eleven characters, counting the two spaces. A model that read three positions now reads eleven. And the cost is worse than it sounds, because of how self-attention works: it compares every position with every other position, building a square grid of scores. Every cell of that grid costs one multiplication.

Think for a moment Three positions gave a grid of \(3 \times 3 = 9\) cells. (Part 1, Section 1.2.3 works one out in full.) How many cells do eleven positions give?

Eleven positions give \(11 \times 11 = 121\) cells, against nine. Figure 4 puts the two plans one above the other at the same scale, so you can see both effects at once.

ONE ROW PER WORD3 positionsIloveLLMsa row for every word the model may ever meet,and no row at all for one it has not metscore grid 3 × 3 = 9 cellsONE ROW PER CHARACTER11 positionsI␣love␣LLMsa row for every character, so no word is ever unknown —and the same three words now fill eleven positionsscore grid 11 × 11 = 121 cellsthe box marked ␣ is a space
Figure 4. The same three words under two plans, at the same cell size: one row per word on top, one row per character below. The plans fail at opposite ends of one trade-off — a table too large to hold, against a score grid of 121 cells instead of 9.

Three words is a small case. A ten-word English sentence runs about fifty characters. Under one row per word the grid is \(10 \times 10 = 100\) cells; under one row per character it is \(50 \times 50 = 2{,}500\). Twenty-five times as much work, for the same sentence.

✗ Common mistake It is tempting to read the two plans as one expensive and one cheap, and to think characters are nearly free, because the table is tiny, but that is wrong. The cost did not go away; it moved. One row per word puts the cost in the lookup table, which is counted once. One row per character puts it in the score grid, and that grid is rebuilt from scratch for every sentence the model reads.
Key takeaway

Try it (2 minutes): Take any sentence of about ten words. Count its words, then count its characters including the spaces. Square both numbers. The second square is the number of cells one self-attention step has to fill if the sentence is cut into characters.

0.3Let the text choose the pieces

We do not have to choose just one side. We can fill the box of stamps with single letters and with parts of words that appear often. This solves both problems at once. Common words get one stamp, so sentences stay short. Rare words get spelled out letter by letter, so no word is ever unknown.

How do we choose these parts? A human does not choose them. We let a computer count them. First, let us define two words:

The computer finds the two pieces that sit next to each other most often, and glues them into one new piece. Then it counts again and glues again. Repeating that is what finds the exact pieces the text repeats most.

This method is called byte-pair encoding, or BPE. Before people used it for language, it was a way to make computer files smaller. In 2016, researchers used it to solve the problem of unknown words in translation. The Transformer model uses byte-pair encoding. You can find both papers in the footer.

The next two sections will walk you through exactly how BPE works:

0.3.1Count the pairs, glue the winner, repeat

You need text to count. A corpus is a large collection of text used for counting. A real corpus has billions of words. Our corpus only has twelve words. We list them below with how many times each word appears. This lets you do the counting yourself.

Two things to know before you read the table:

wordtimeswordtimes
love11grand6
LLM12grands3
LLMs3modèle12
aime6modèles2
les14I4
J2’3

Here is the loop. Four steps, and the only arithmetic in it is addition.

Notice one very important thing about these four steps: the vocabulary only grows. The original eighteen letters are still there at the end. This is why the model can always spell out an unknown word letter by letter.

0.3.2Four merges, counted by hand

Figure 5 runs the loop four times over that corpus. Each stage shows three things: every corpus word cut into the pieces it has now, the neighbouring pairs ranked by total, and the size of the vocabulary. Press play, or read the four stages below, which give the same numbers in words.

STARTevery word cut into charactersCORPUSTIMESCUT INTOlove11loveLLM12LLMLLMs3LLMsaime6aimeles14lesgrand6grandgrands3grandsmodèle12modèlemodèles2modèlesI, J and ’ are one character each, so they make no pairs.TOP 8 ADJACENT PAIRS, MOST FREQUENT FIRSTl + e→le28e + s→es16L + L→LL15L + M→LM15m + o→mo14o + d→od14d + è→dè14è + l→èl14one pair is ahead of every other, so no rule is needed here.The winner becomes merge 1.VOCABULARY18 tokensAFTER MERGE 1merged le, count 28CORPUSTIMESCUT INTOlove11loveLLM12LLMLLMs3LLMsaime6aimeles14lesgrand6grandgrands3grandsmodèle12modèlemodèles2modèlesI, J and ’ are one character each, so they make no pairs.TOP 8 ADJACENT PAIRS, MOST FREQUENT FIRSTle + s→les16L + L→LL15L + M→LM15m + o→mo14o + d→od14d + è→dè14è + le→èle14l + o→lo11one pair is ahead of every other, so no rule is needed here.The winner becomes merge 2.VOCABULARY19 tokensAFTER MERGE 2merged les, count 16CORPUSTIMESCUT INTOlove11loveLLM12LLMLLMs3LLMsaime6aimeles14lesgrand6grandgrands3grandsmodèle12modèlemodèles2modèlesI, J and ’ are one character each, so they make no pairs.TOP 8 ADJACENT PAIRS, MOST FREQUENT FIRSTL + L→LL15L + M→LM15m + o→mo14o + d→od14d + è→dè14è + le→èle12l + o→lo11o + v→ov112 pairs tie at 15, so the rule decides: take the one that appears first.The winner becomes merge 3.VOCABULARY20 tokensAFTER MERGE 3merged LL, count 15CORPUSTIMESCUT INTOlove11loveLLM12LLMLLMs3LLMsaime6aimeles14lesgrand6grandgrands3grandsmodèle12modèlemodèles2modèlesI, J and ’ are one character each, so they make no pairs.TOP 8 ADJACENT PAIRS, MOST FREQUENT FIRSTLL + M→LLM15m + o→mo14o + d→od14d + è→dè14è + le→èle12l + o→lo11o + v→ov11v + e→ve11one pair is ahead of every other, so no rule is needed here.The winner becomes merge 4.VOCABULARY21 tokensAFTER MERGE 4merged LLM, count 15CORPUSTIMESCUT INTOlove11loveLLM12LLMLLMs3LLMsaime6aimeles14lesgrand6grandgrands3grandsmodèle12modèlemodèles2modèlesI, J and ’ are one character each, so they make no pairs.TOP 8 ADJACENT PAIRS, MOST FREQUENT FIRSTm + o→mo14o + d→od14d + è→dè14è + le→èle12l + o→lo11o + v→ov11v + e→ve11g + r→gr93 pairs tie at 14, so the rule decides: take the one that appears first.The winner becomes merge 5.VOCABULARY22 tokens
Figure 5. The merge loop, one stage per merge: corpus words and their current pieces on the left, pairs ranked by total in the middle, vocabulary count on the right. Watch the modèles row, where the token les ends up serving as the last three letters of a longer word.

Merge 1: l and e (total 28). The pair l e appears in three words: les (14 times), modèle (12 times), and modèles (2 times). That adds up to 28. Nothing else comes close. The next best is e s with 16. So, we glue l and e into one token: le. Our vocabulary grows from 18 to 19 tokens.

Think for a moment Now that le is one piece, the pair le s exists where e s used to be. Which two words hold this new pair, and what is its total count?

Merge 2: le and s (total 16). The answer to the question above is 16. The pair le s appears in les (14 times) and modèles (2 times). The next best pairs are L L and L M, both with a count of 15. Since 16 beats 15, we glue le and s to make the token les. Our vocabulary grows to 20 tokens. Notice that after just two steps, the computer has discovered a complete French word (les), without knowing any French!

Merges 3 and 4 continue the exact same way. The table below shows the winning pair for each step.

mergewinnertotalfromvocabulary
3L + L → LL15 LLM 12, LLMs 320 → 21
4LL + M → LLM15 LLM 12, LLMs 321 → 22

At Merge 3, we have a tie. Both L L and L M appear 15 times. Which one should we pick? Ties happen constantly in text, so we need a rule.

We cannot just pick randomly. The rule must be fixed. If the model splits the same sentence two different ways on two different days, it will fail. You must pick a tie-breaking rule and consistently stick to it.

Our rule here is: the pair that appears earliest in our corpus list wins. Reading left-to-right, L L appears before L M in the word LLM. So, L L wins Merge 3.

In Merge 4, the new pair LL M also has a count of 15, so it easily wins, giving us the token LLM. The remaining pair LLM s only appears 3 times, which is too low to win right now.

Key takeaway

Try it (3 minutes): Merges 5, 6 and 7 are mo, mod and modè. Work out the total each of them had when it won, using only the corpus table. All three should come to 14, and the two words that supply it are the same two every time.

0.4The two sentences, cut up

In Section 0.3, we watched the computer build the first four tokens: le, les, LL, and LLM. What happens if we keep letting it run?

We let the loop run eighteen times over our mini-corpus, and then we stop it. The loop could keep going: the pair LLM s is still sitting there with a count of 3. But eighteen merges is enough for this page, and by then nine of our twelve words are a single piece each. Here is the final list of all eighteen merges. The numbers 1 to 18 represent the exact step when the piece was created:

Steps 1–6Steps 7–12Steps 13–18
Step 1: leStep 7: modèStep 13: gra
Step 2: lesStep 8: modèleStep 14: gran
Step 3: LLStep 9: loStep 15: grand
Step 4: LLMStep 10: lovStep 16: ai
Step 5: moStep 11: loveStep 17: aim
Step 6: modStep 12: grStep 18: aime

Are these 18 pieces the only tokens in our vocabulary? No. Remember that BPE never deletes old tokens. We started with 18 single characters (like a, d, e). We just added these 18 new merged pieces. So, our final vocabulary has exactly 36 tokens. Our token list is now finished and ready to use.

0.4.1The merge list is the tokeniser

We are done counting. The counting loop (Section 0.3) was only used to build the merge list. You only do this once, before you even start training the model.

Now, we switch to a different job: using the list. The model does this every time someone types a sentence.

Phase 1: Building the listPhase 2: Using the list
When?Only onceEvery time a user types
InputA massive corpus of textOne single sentence
JobCount pairs and find the winnersJust follow the saved merge list
Counting?YesNo counting at all

To use the list on a new sentence, we follow three steps:

The order guarantees the exact same result every time. LLMs meets merge 3 before merge 4. This means everyone always gets LLM and s. Figure 6 runs the list over I love LLMs, one merge per stage. It starts from nine characters rather than the eleven Section 0.2 counted, because step one has already taken the two spaces out.

STARTcut at the spaces first, so the two spaces are goneIloveLLMsIloveLLMs9 charactersTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysMERGE 3L + L becomes LLIloveLLMsIloveLLMs8 leftTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysMERGE 4LL + M becomes LLMIloveLLMsIloveLLMs7 leftTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysMERGE 9l + o becomes loIloveLLMsIloveLLMs6 leftTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysMERGE 10lo + v becomes lovIloveLLMsIloveLLMs5 leftTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysMERGE 11lov + e becomes loveIloveLLMsIloveLLMs4 leftTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysDONEfour tokens, for three wordsIloveLLMsIloveLLMs4 tokensTHE MERGES THAT TOUCH THESE WORDS3.  LL4.  LLM9.  lo10.  lov11.  loveapplied in this order, alwaysnothing is counted here · the list was built in Section 0.3 and is only being read
Figure 6. The merge list applied to I love LLMs, one merge per stage, with the five merges that touch these words listed on the right. Nine characters become four tokens, and the only thing deciding the outcome is the order the merges are read in.

Here is what the list does to all nine of them.

cut at spacesafter the 18 mergestokens
II1
lovelove1
LLMsLLM s2
JJ1
’’1
aimeaime1
lesles1
grandsgrand s2
modèlesmodè les2

So I love LLMs is four tokens for three words. J’aime les grands modèles is eight tokens for four words: J ’ aime les grand s modè les.

Why did LLMs split into two tokens instead of staying as one? Look at the counts. The pair LLM s appears only 3 times in our corpus, because the only word that contains that pair is LLMs, and LLMs appears 3 times. Each round of BPE picks the pair with the highest count. So a pair with count 3 has to wait until every pair with count 4 or higher has already been merged. We only ran 18 rounds, and there were still pairs with count 6 or higher being merged at the end. The pair LLM s never reached the front of the line, so it was never glued together.

grands split for the same reason, and at the same total of 3. So the plural s is a token on its own, shared by both words. That is the fix for the third cost in Section 0.1, where walk, walks, walked and walking each wanted a row.

✗ Common mistake It is tempting to think BPE learns grammar, but that is wrong. It only counts characters: grand and s come apart because that pair is frequent, not because s is a plural. BPE finds grammar boundaries by accident, because grammar rules naturally create frequent spelling patterns.

0.4.2A character the corpus never held

Earlier, we promised that the model would never see an unknown word. But our alphabet only has eighteen characters. If you type a German ü or a Chinese 字, the model will not recognise them. The unknown-word problem has become an unknown-character problem.

To fix this, we start one level below characters: with bytes. Computers store all text as bytes (numbers from 0 to 255). We use a system called UTF-8, which turns every character into one to four bytes. Plain English letters take one byte each, è takes two, and the curly apostrophe in J’aime takes three. Our French sentence is 25 characters and 28 bytes.

Start the alphabet at all 256 byte values instead of at the characters of a corpus, and the problem cannot recur. Every possible text, in every language, in every alphabet, is some sequence of bytes, and all 256 of them are already in the vocabulary. That choice is called byte-level byte-pair encoding. GPT-2 uses it.

?  The question è is two bytes, so byte-level BPE starts it as two separate pieces. Does that break the character in half? Yes, at the start, and only at the start. Those two bytes only ever appear together, because together they are that character. Their pair count is therefore the count of è itself, which is high, so the loop glues them back into one piece in its first few rounds.

What changes in the loop? The counting does not change. Step one cuts into bytes rather than characters. Steps two, three and four are word for word the same: count the neighbouring pairs, glue the most frequent, repeat. What changes for you is the size of the arithmetic. modèles starts as eight pieces rather than seven, and the curly apostrophe in J’aime starts as three. Counting the pairs stops being something you can do with a pencil. That is why this page counts characters and not bytes.

Key takeaway

Try it (2 minutes): Tokenise les modèles using only the merge list above. You should get three tokens, and one of them should appear twice.

0.5A token id is a row number

Our vocabulary has 36 tokens. Number them from 0 to 35. This number is called a token id. It is just the token's place in the list. It has no special meaning, and nearby numbers do not mean the words are similar.

Order the 18 characters first, in the order Section 0.3.1 listed them, which takes ids 0 to 17. Then the 18 merges, in the order they were made, so merge \(k\) takes id \(17 + k\). The four tokens of I love LLMs then get ids 0, 28, 21 and 14, and you can check every one against the alphabet in Section 0.3.1 and the merge list in Section 0.4. Counting from 0, I is character number 0 and s is character number 14, so those are their ids. LLM is merge 4, so its id is \(17 + 4 = 21\), and love is merge 11, so its id is \(17 + 11 = 28\). Figure 7 lays the whole vocabulary out as one numbered strip.

THE VOCABULARY, NUMBERED FROM 018 characters first, then the 18 merges in the order they were madeI0JLM3ade6gil9mno12rs14v15è’le18lesLLLLM21momodmodè24modèlelolov27love28grgra30grangrandai33aimaime← characters, ids 0 to 17merges, ids 18 to 35 →I love LLMsIid 0loveid 28LLMid 21sid 14← each id is a place in the strip above, nothing more
Figure 7. The vocabulary as one numbered strip, characters first and merges after, with the four tokens of I love LLMs pointing at their ids. An id is a place in this strip and carries no other meaning.

0.5.1The three markers

We add three more tokens by hand. The counting process would never find them because they are not real text. The model uses them as special flags to control how it works. The first two are for the half of the model that writes the French sentence out, one word at a time — the decoder. The third has nothing to do with writing at all.

They are tokens like any other. Each gets a row and each gets an id. The model reads them and can write them, exactly as it reads and writes love. The angle brackets are a writing convention, so that you can tell a marker from a word on the page, and the only thing that makes a marker special is that a person reading the output discards it.

Add the three, and this part's vocabulary is 39 tokens.

0.5.2The largest table in the model

The vocabulary is a list of token strings. The model cannot do arithmetic on strings, so it needs a table that turns each token into a row of numbers. That table is called \(\mathbf{E}\). It has one row per token and \(d_{\text{model}}\) numbers in each row. Here are four of its 39 rows, at the toy width \(d_{\text{model}} = 4\) that Parts 1 to 8 work with. The ids are the ones you worked out above; the four numbers in each row are made up, because nothing on this page computes them.

idtoken\(\mathbf{E}\) row
0I0.12−0.340.560.09
14s−0.190.510.07−0.63
21LLM0.450.67−0.220.30
28love−0.710.030.880.41

The numbers in each row are learned during training. Part 9, Section 9.9.1 moves one row of \(\mathbf{E}\) by hand, so you can see what learned means. Before training they are random; after training they carry meaning, but this part does not need to know what they are. All this part does is look them up.

Looking a token up means: take the token's id, go to that row of \(\mathbf{E}\), and copy the row. If the sentence is I love LLMs (four tokens with ids 0, 28, 21, 14 in our full vocabulary), you copy four rows and stack them top to bottom. That stack is \(\mathbf{X}\), the matrix Part 1 starts from. Figure 8 follows the whole path.

1 CHARACTERS11 of themthe box marked ␣ is a spaceI␣love␣LLMsfirst cut at the spaces: 9 pieces leftthen apply the 18 merges in order2 TOKENS4 of them, for 3 wordsIloveLLMsLLMs is two tokens: the word and the plural slook each token up in the vocabulary3 TOKEN IDSa row number, nothing more0282114take that row out of E4 ROWS OF Xone row per token, d_model numbers wideI…love…LLM…s…this stack of rows is the matrix Xthat Part 1 starts fromFour rows, because this sentence is four tokens. Parts 1 to 8 give each word one row and so work with three.
Figure 8. I love LLMs in four bands: eleven characters, nine pieces once the two spaces are cut away, four tokens after the merge list has run, four ids, then the rows those ids pull out of \(\mathbf{E}\). Every arrow here is a lookup or a substitution — nothing on this page is learned or multiplied.
! One simplification, declared here Figure 8 ends with four rows, because I love LLMs is four tokens under the vocabulary this part built. Parts 1 to 8 give each word one row, so they work with three rows, and their \(\mathbf{E}\) has ten rows rather than 39 — seven words between the two sentences, plus the three markers. That is a simplification of the same kind as leaving the positional encoding out of the worked arithmetic, and it keeps the matrices small enough to multiply on paper. Everything those parts do would work the same way with four rows and 39 tokens; every step would just be longer to write out.
A massive table — Part 5, Section 5.7 counts 18,902,016 numbers in the six blocks of the encoder stack, and \(\mathbf{E}\) holds 18,944,000. The table that turns tokens into numbers is as heavy as the whole stack of blocks that then processes them.

Now the size. The paper's base model has \(d_{\text{model}} = 512\), and its English-to-German experiment shares a vocabulary of about 37,000 tokens between the two languages. So \(\mathbf{E}\) holds \(512 \times 37{,}000 = 18{,}944{,}000\) numbers. That is the largest single table the model actually has — smaller than the 51,200,000 Section 0.1 rejected, and bigger than any other single matrix in the model by a factor of about eighteen.

Byte-pair encoding brought that figure down to 18,944,000 for two languages at once, and removed the unknown-word problem while doing it. Both of those come from one decision. 37,000 is not a count of anything in either language; it is the number of merges the paper stopped at.

Key takeaway

Try it (2 minutes): The vocabulary here is 36 tokens plus 3 markers. Work out how many numbers \(\mathbf{E}\) would hold at \(d_{\text{model}} = 512\), then say how many times larger the paper's table is. You should get 19,968, and the paper's is about 949 times that.

0.6Turning ids back into text

When the model writes its translation, each step produces one probability for every token in the vocabulary. Part 6, Section 6.4.2 works one of those out, and Part 10 is about how a token gets chosen from them. Picking one gives a token id. To show the answer to a person, we have to turn those ids back into readable text. The first step is easy: we just look up each id in the vocabulary to get its string. For example, the ids 0, 28, 21, 14 give us the pieces I, love, LLM, s.

But when we try to put those pieces back into a sentence, we hit a problem. Where did the spaces go? Pre-tokenising threw them away at the very beginning of the process, and nothing has held them since. If you just join the four pieces, you get IloveLLMs.

Think for a moment How do we remember that a space goes before love but not before s? The model only sees token ids, so the space must be inside the token itself.

We attach the space to the front of the token. So, ␣love becomes the token, not love. This makes joining easy: just stick the pieces together. Put the token strings end to end, and the spaces are already in the right places. Figure 9 joins the four tokens both ways. Those space-carrying tokens would take ids of their own; the vocabulary this part built does not have them.

JOIN THE FOUR TOKENS BACK TOGETHERIloveLLMs→IloveLLMs✗pre-tokenising threw the spaces away at the first step, and nothing has held them sinceI␣love␣LLMs→I love LLMs✓the space rides at the front of the token that follows it, so joining is plain concatenation␣love is a different token from love
Figure 9. The same four tokens joined twice: without the space the words run together, and with the space carried at the front of each token they come back apart. The space has to live inside a token, because a token is the only thing the model ever sees.

One important consequence for Part 6: when the model generates text, it generates one token at a time, not one word at a time. For example, modèles is split into modè and les, so the model takes two steps to write it. Part 6 will pretend the model writes whole words, but that is just a simplification to keep the arithmetic small.

Key takeaway

NextThe matrix Part 1 starts from

You now know how to take a sentence, chop it into tokens using a merge list, look up each token's row in \(\mathbf{E}\), and stack those rows together. That stack of rows is \(\mathbf{X}\). In Part 1, Section 1.2.1, \(\mathbf{X}\) is simply taken as given — but now you know exactly where it comes from.

One reading rule for everything that follows. Whenever Parts 1 to 8 say word, they actually mean token. They use word because their toy examples pretend every word is exactly one token, which keeps the arithmetic small. But in a real model, a single word can be split into many tokens. Just remember: every row is one token, every position is one token, and the vocabulary is just a list of tokens.

Part 1 is self-attention: where Q, K and V come from, and the whole formula worked by hand on I love LLMs. It also says why the scores are divided by \(\sqrt{d_k}\). Part 1, Section 1.2.1 works with the same matrix Section 0.5 has just built, and takes it as given. Section 1.1 comes first, and is about why attention exists at all.