Understanding the techniques behind every chatbot, with open-weights models you can download, run and study yourself.
The diagram inside every chatbot. By the end of this you can read every box.
You will find out what is actually happening when one answers you, and why it behaves the way it does.
Tokens are the bill, the context window is the memory, temperature is the dice. All three are explained here.
Every number on these slides is sourced, and the last slide has the commands to reproduce every demo yourself.
Make your guess before you hit reveal. You are about to do, by hand, the only thing a language model ever does.
Real numbers: gemma3:4b, an open-weights model (4-bit, as Ollama ships it), recorded offline in raw next-token mode, September 2026. The top 5 of the 262,144 tokens it knows.
Every idea, no equations. One sentence carried through all nine stations.
ChatGPT, Gemini, Claude, DeepSeek: a chat window wrapped around an LLM that was also trained to follow instructions and hold a conversation.
A program trained on a huge pile of text to do one thing: predict the next token. It never looks anything up. Large means billions of learned numbers; gemma3:4b, used in the demos here, has 4 billion.
Text becomes tokens, tokens become lists of numbers, and the same block is stacked again and again. It reads chunks, not letters, and only what fits in its window.
Every word looks at every other word, all at once, and picks the ones that matter. That is how it finds animal.
They threw out the step-by-step reading of older models and kept only attention. That one idea now runs every chatbot you use.
"We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely."
A better translator, built only from attention.
Open weights is not open source. You can download and run these models, but the training data and code mostly stay hidden. OLMo publishes all three.
Know why it hallucinates, forgets and miscounts, and work around it on purpose.
Tokens are the bill, the context window is the memory, temperature is the dice.
Llama, DeepSeek, Qwen, Gemma: the same block with different settings. Try one tonight.
Open models run on a laptop. Free software needs people who understand what is inside.
Every demo here runs on your own machine.
We use Ollama, a free tool that downloads an open model and runs it offline. Install it, pull one small model, and you can reproduce every demo in this deck. Where a demo can be re-run, the command is printed next to it.
ollama pull gemma3:1b # 815 MB, runs on a laptop
Ollama is a free, open-source tool that downloads an open-weights model and runs it entirely offline. Nothing leaves your machine, and there is no account or API key.
ollama pull gemma3:1b # 815 MB ollama pull gemma3:4b # 3.3 GB, optional
gemma3:1b is Google's smallest open model. It is deliberately small, which is exactly why it shows the rough edges so clearly.
ollama run gemma3:1b "Why is the sky blue?"
That is the whole setup. From here, wherever a demo can be reproduced, the exact command sits on the slide next to it, marked run it yourself.
A fair warning. A 1-billion-parameter model on a laptop is far weaker than the chatbots you are used to. That is the point: the failures we are about to explain are easiest to see on a small model, and the mechanism behind them is identical in the large ones.
Word 7 has to wait for word 6. A graphics card with 10,000 cores sits mostly idle.
Information from early words fades with every hop. Long sentences lose their beginning.
Maximum path length between any two words: O(1) for self-attention, O(n) for recurrence. That one column is the paper's whole argument.
The price: O(n²·d) work per layer. Double the text, quadruple the attention work.
It gets three chunks and never the letters inside them, so counting the r's means guessing. gemma3:1b answered ten. Bigger models often get it right now, by guessing better, not by seeing letters.
ollama run gemma3:1b "How many r's in strawberry?"Tokens are what you pay for. API prices and context limits are all counted in tokens.
Start from 256 raw bytes. Merge the most frequent neighbouring pair. Repeat about 50,000 times. GPT-2 ends with 50,257 tokens; Llama 3 uses 128,256 and Gemma 3 about 262,000.
Kannada "ಕನ್ನಡ ಭಾಷೆ" costs GPT-2 27 tokens for 10 characters.
cos(a, b) = a·b / (|a| |b|)
Real vectors: 768 numbers per token in GPT-2, 4,096 in Llama 3 8B. Position in 2017: add sine waves. In 2026: RoPE rotates Q and K by position.
The numbers say what a word is, not where it sits. So every token also gets a position stamp.
You just did attention yourself: you read the whole sentence, worked out which word it points to, and ignored the rest. A Transformer does this for every word, in every layer. Click any word to see its pattern.
The weights shown are illustrative, shaped like the heads in the paper's Figures 3 to 5. A trained model computes a set like this for every word at once.
Each arc is one attention weight: how much it borrows from that word. The weights for one word always add up to 100%. These are illustrative, shaped like the real heads in the paper's Figures 3 to 5, and a real model computes a set like this for every word at once.
Your question (Query) is compared with every book's label (Key), and the best matches hand over their contents (Value). Not one book: a blend, weighted by how well each one matched.
q = x·WQ k = x·WK v = x·WV
Why not compare the raw word vectors? "it" is not similar to "animal", it is relevant to it. WQ and WK learn relevance. Nobody sets them by hand: training does.
We followed one word, it, through one row of this grid. The model fills every row at the same moment: each word asks its own question, and each word is also a Key and a Value for all the others.
A 1,000-token chat means a million scores, in every head of every layer. Nothing waits for anything else, which is exactly the work a graphics card is built for.
This grid is softmax(QKT/√dk): one row per query, one softmax per row, all rows in one matrix multiply. In a chat model (a decoder) the hatched half is masked: a word may only look back. So in GPT-style models "it" cannot see "tired"; the link is made later, when "tired" looks back at "animal".
"it" asks a question, every word answers with a match score, softmax turns the scores into shares of 100%, and "it" takes that mix of their meanings. Those shares are the attention.
Attention(Q,K,V) = softmax( QKT / √dk ) · V
A dot product over dk numbers grows like √dk. Without the division, big heads give huge scores, softmax bets everything on one word, and the gradient dies. Try it: turn scaling off and drag dk.
Illustrative roles, modelled on heads found in trained models (Clark et al., 2019). Many real heads are messier or redundant (Michel et al., 2019).
Each head has its own Query, Key and Value, so each one learns to look for a different kind of link in the same sentence. Nobody assigns these jobs: they appear during training. These four are illustrative. Some real heads are this clean; many are messier.
All heads run at the same moment. Their answers are joined side by side, then mixed into one update for each word. The 2017 paper used 8 heads. Llama 3 8B has 32 in each of its 32 layers: 1,024 heads.
X is the sentence, one row per token. Each head has its own WQ, WK, WV: Q = XWQ, QKT is an n × n grid, softmax each row, times V. Then Concat(head1…head8)·WO mixes the heads. 2017 base: 512 numbers → 8 heads of 64. Llama 3 8B: 32 query heads share 8 K/V sets.
Words share notes with each other. This is the only place they talk.
Each word goes back to its own desk and thinks about what it heard. About two thirds of the weights live here.
Each layer writes a small edit on top of what came before, so nothing is lost on the way up.
6 blocks in 2017, 32 in Llama 3 8B, 61 in DeepSeek-V3. The same block, and the diagram on the cover.
FFN(x) = W2 · max(0, W1x + b1) + b2
512 → 2048 → 512 in the base model. Each sublayer is wrapped as x + Sublayer(x), then LayerNorm ("Add & Norm").
Three wirings: encoder-only (BERT) · decoder-only (Llama, Qwen, DeepSeek, Gemma, GPT) · encoder-decoder (the 2017 original, T5, Whisper).
Figure 1 of the paper. You can now name every box in it. Vaswani et al., 2017, reproduced with the paper's stated permission.
gemma3:4b's real top 8 for this sentence, rescaled to add up to 100%. Rolls are seeded, so the demo repeats.
Every station turned each word into numbers full of context. The model's real job, the game from slide 4, is to guess the word that comes next.
So hide tired. The last word, too, now carries the whole sentence. With tired gone, "it" could be the animal or the street again, so the guesses split.
Low: it always takes the top word, safe but repetitive. High: long shots win and nonsense creeps in. It is not a creativity dial.
pi = ezi/T / Σj ezj/T
While writing, token n may only look back at tokens 1 to n: the causal mask. That is also why training can predict every position of a sentence at once.
Each new word depends on every word before it, including the model's own last word. Word 5 cannot be guessed until word 4 exists.
The past never changes, so its keys and values are kept. Each new pass only computes the newest token.
It is everything the model can read at once: your messages, its replies, all of it, counted in tokens. When the chat outgrows it, the oldest part falls out. No error, no warning.
Each lab's model docs, 24 Sep 2026. 1M tokens ≈ 2,000 to 3,000 pages.
Attention compares every token with every token: O(n²). Double the window and that work quadruples, and the KV cache grows with every token.
"Lost in the middle" (Liu et al., 2023): facts buried mid-context get used less. A tendency, strongest in smaller and older models.
A plausible abstract is the most likely continuation. Truth was never part of the objective. Bigger models with search and tools do this far less: same mechanism, lower rate.
ollama run gemma3:1b "Summarise the 2019 paper 'Attention Fields in Sparse Halting Networks' in three sentences."Training minimises next-token error. Nothing in that rewards "I don't know". OpenAI's 2025 paper Why Language Models Hallucinate argues training and benchmarks reward confident guessing. Fixes: retrieval, tools, training to abstain.
"...introduces a novel method for approximating the outputs of sparse halting networks, which are used in reinforcement learning … achieving impressive results in tasks like Atari games."
"The paper introduces Halting Networks, a novel architecture for predicting whether a program will halt, using a sparse, learned attention mechanism over the execution trace."
Real outputs, same prompt, recorded offline. Two models, two completely different inventions. If they were looking something up, they would agree.
Totals as each developer states them on its model card. Closed labs (GPT-4 onward, Claude, Gemini) no longer publish sizes, so they cannot be plotted.
Attention, then MLP, with a residual stream, stacked. The same block you just built.
Many small expert MLPs, and a router picks a few per token. Kimi K3 (2026, open weights) has 2.8 trillion weights but uses only 104 billion for each token.
RoPE: rotate Q and K by position instead of adding sine waves.
KV cache: never recompute the past.
GQA: many query heads share one set of K and V.
FlashAttention: same maths, tiled so the n × n grid never hits memory.
Mixture of experts: many MLPs, a router picks a few.
Read any model paper with four questions: how do tokens talk, what is the MLP, how is position encoded, what is it trained to predict?
Images · 16×16 patches (ViT)
Audio · slices of sound (Whisper)
Proteins · amino acids (AlphaFold)
Code · a variable finds where it was declared
1. Install Ollama (free, MIT licence) and run ollama run gemma3:1b
2. Also try llama3.2, qwen3, deepseek-r1: all open weights
3. Ask one about a paper that does not exist
4. Then Karpathy's "Let's build GPT": 2 hours, about 300 lines
Softmax, worked out · how Q, K and V are learned · the full command sheet · thousandmiles.ai
First given as a live session for engineering students new to LLMs, then rebuilt to read on its own.
softmax(z)i = ezi / Σj ezj
Exponentiating makes every number positive. Dividing by the sum makes them add to 1. So any list of scores becomes a valid set of weights.
It leans towards the winner. A gap of 2 in the scores becomes a factor of about 7.4 in the weights, so attention stays selective instead of a bland average.
It is smooth everywhere, so gradients flow back into WQ and WK. A hard "pick the max" would give no gradient at all. Real code also subtracts the max score first so ez never overflows.
At the start it is noise. WQ, WK, WV are random, queries match keys at random, and the guesses are gibberish.
One thing is measured. Show it "The animal didn't cross the street because it was too", ask for the next word, compare with the real one. One number: how wrong.
The error flows backwards. From the loss, through the output, the blend, the softmax and the dot products, into WQ and WK. Every step is differentiable, which is why attention is built from dot products and a softmax.
A small nudge. "This guess would have been better if 'it' had looked more at 'animal'." So the matrices shift to line that query and key up a little. Repeat over billions of tokens.
Grammar falls out as a side effect. Pronoun resolution was never a goal. It is simply what helps predict the next word, and the same pressure produces heads nobody has a name for.
ollama --version # needs 0.12.11+ for logprobs ollama pull gemma3:1b # ~800 MB ollama pull gemma3:4b # ~3.3 GB, the numbers in this deck
curl localhost:11434/api/generate -d '{
"model": "gemma3:4b", "raw": true, "stream": false,
"prompt": "The capital of France is",
"logprobs": true, "top_logprobs": 5,
"options": {"num_predict": 1}
}' # probability = exp(logprob)
ollama run gemma3:1b /set parameter temperature 0 # ask 3 times: identical /set parameter temperature 1.5 # ask 3 times: all different
/set parameter num_ctx 256 # "The secret code is BANANA." + 300 tokens of filler # then ask for the code: it cannot say /set parameter num_ctx 4096 # now it can
ollama run gemma3:1b "Summarise the 2019 paper \"Attention Fields in Sparse Halting Networks\" in three sentences."
Pre-warm with "keep_alive": -1 or the first call stalls while the model loads. Use "raw": true for pure next-token numbers, or the chat template changes them.