A tutorial series  ·  Eleven parts

The Transformer, Number by Number

Every step of the transformer worked out by hand on one three-word sentence. No step is skipped, and every number in between is printed on the page.

0192837465019283 8475639201847563 2938475610293847 5748392019574839 1029384756102938 6574839201657483 3847561029384756 9283746501928374
The series mark. Optimus Prime is the other kind of transformer, and this one is built number by number. Go close enough to any plate of his armour and there is nothing there but digits. The model these eleven pages take apart is made the same way.
ENCODER · runs onceI love LLMs+ positional encodingat the bottom, oncesix encoderblocksPart 5C · three rowsDECODER · runs once per word<start> J’aime les …+ positional encodingat the bottom, every runsix decoder blockseach: self-attention,cross-attention,feed-forwardSection 7.5side doorK and V, every runoutput head · Part 6the next wordwritten down, then fed backin at the bottom next runThe encoder runs once and its three rows never change. The decoder runs again for every word,one row longer each time, and reads those same three rows through the side door on every run.
The architecture. An encoder that runs once, and a decoder that runs again for every word it writes. The side door carries the encoder’s rows across on every run. The paper draws this as its Figure 1. Here it is the last figure of Part 7, reached one piece at a time.

01What this is

A transformer is the architecture behind today’s large language models, and behind machine translation before them. Most explanations of it stop at the diagram: a box marked attention, an arrow into a box marked feed-forward. You reach the end and you still cannot say what number comes out of the first box.

This series does the arithmetic instead. One sentence, I love LLMs, becomes a matrix of three rows with four numbers in each row. That matrix then goes through every step of the model in Attention Is All You Need (Vaswani and others, 2017). Every value in between is printed, and you can check any of them with a pencil.

Who this is for

Someone who has met matrices before and wants to know what a transformer actually computes. A student taking a first course on deep learning. An engineer who uses these models and would like to see inside one. A reader who has tried the paper and stopped at Section 3.

You are assumed to be new to this subject, not new to everything. Terms specific to transformers are defined where they first appear, and so is every symbol.

02How to read it

Start at Part 1 and read through to Part 10. Each part from Part 2 onward opens with a short paragraph naming what it assumes, and which earlier section built each piece. You can tell in one look whether you are ready.

Part 0 sits outside that order, and builds the matrix Part 1 starts from. Read it first, or leave it until after Part 8. Nothing in Parts 1 to 8 depends on having read it.

You need matrix multiplication, transpose, and the dot product. Everything past that is built on the page: softmax, layer normalisation, residual connections, the two masks, the loss, and the gradient.

03The eleven parts

Eleven characters become tokens, tokens become id numbers, and id numbers become rows. Byte-pair merges counted by hand, the three markers, and the embedding table.
What recurrence costs, what Query, Key and Value each are, and the whole attention formula worked out on three words.
One head gives every token a single budget to spend. Why that is the wrong shape for language, what splitting splits, and how many heads to use.
Shuffle the sentence and the output rows only move with it. The proof that order is lost, why no later layer can put it back, and why a plain counter fails.
The sine and cosine formula worked by hand, where it is added, why it lets a model reason about distance, and the rotary version that current models use instead.
Attention alone cannot answer a yes-or-no question about a row. The three things wrapped around it — the residual addition, layer normalisation, the feed-forward network — and the finished encoder block.
A row leaving the stack is not a word. The output head that turns one row into a word, the loop that feeds the word back in, and the two markers the loop needs.
Cross-attention is the step that carries the encoder’s rows into the decoder. One such step worked by hand, the three sublayers of a decoder block, and the whole transformer in one figure.
On the second run the multiplication produces one score the rules forbid. The causal mask that stops it counting, and the padding mask for sentences of different lengths sharing one batch.
Every weight so far was written down by hand. Training is the correction: a sentence somebody already wrote becomes a number, the number becomes a direction, and the direction changes every weight.
Ten probabilities have to become one word. Greedy choosing, sampling, temperature, top-k and top-p, beam search, when to stop, and what the KV cache saves.