The series mark. Optimus Prime is the other kind of transformer, and this one is built
number by number. Go close enough to any plate of his armour and there is nothing there but digits. The
model these eleven pages take apart is made the same way.
The architecture. An encoder that runs once, and a decoder that runs again for every
word it writes. The side door carries the encoder’s rows across on every run. The paper draws this
as its Figure 1. Here it is the last figure of Part 7, reached one piece at a time.
01What this is
A transformer is the architecture behind today’s large language models, and behind machine translation
before them. Most explanations of it stop at the diagram: a box marked attention, an arrow into a box
marked feed-forward. You reach the end and you still cannot say what number comes out of the first
box.
This series does the arithmetic instead. One sentence, I love LLMs, becomes a matrix of three rows
with four numbers in each row. That matrix then goes through every step of the model in Attention Is
All You Need (Vaswani and others, 2017). Every value in between is printed, and you can check any of them
with a pencil.
Who this is for
Someone who has met matrices before and wants to know what a transformer actually computes. A student
taking a first course on deep learning. An engineer who uses these models and would like to see inside
one. A reader who has tried the paper and stopped at Section 3.
You are assumed to be new to this subject, not new to everything. Terms specific to transformers are
defined where they first appear, and so is every symbol.
02How to read it
Start at Part 1 and read through to Part 10. Each part from Part 2 onward opens with a short paragraph naming
what it assumes, and which earlier section built each piece. You can tell in one look whether you are
ready.
Part 0 sits outside that order, and builds the matrix Part 1 starts from. Read it first, or leave it until
after Part 8. Nothing in Parts 1 to 8 depends on having read it.
You need matrix multiplication, transpose, and the dot product. Everything past that is built on the page:
softmax, layer normalisation, residual connections, the two masks, the loss, and the gradient.
Eleven characters become tokens, tokens become id numbers, and id numbers become rows.
Byte-pair merges counted by hand, the three markers, and the embedding table.
Shuffle the sentence and the output rows only move with it. The proof that order is lost,
why no later layer can put it back, and why a plain counter fails.
The sine and cosine formula worked by hand, where it is added, why it lets a model reason
about distance, and the rotary version that current models use instead.
Attention alone cannot answer a yes-or-no question about a row. The three things wrapped
around it — the residual addition, layer normalisation, the feed-forward network — and the
finished encoder block.
A row leaving the stack is not a word. The output head that turns one row into a word, the
loop that feeds the word back in, and the two markers the loop needs.
Cross-attention is the step that carries the encoder’s rows into the decoder. One
such step worked by hand, the three sublayers of a decoder block, and the whole transformer in one figure.
On the second run the multiplication produces one score the rules forbid. The causal mask
that stops it counting, and the padding mask for sentences of different lengths sharing one batch.
Every weight so far was written down by hand. Training is the correction: a sentence
somebody already wrote becomes a number, the number becomes a direction, and the direction changes every
weight.
Ten probabilities have to become one word. Greedy choosing, sampling, temperature, top-k
and top-p, beam search, when to stop, and what the KV cache saves.