thousandmiles.ai

Open the black box: how LLMs actually work

Understanding the techniques behind every chatbot, with open-weights models you can download, run and study yourself.

Figure 1 of Attention Is All You Need: the Transformer architecture
Figure 1 · Vaswani et al., 2017 · the Transformer

The diagram inside every chatbot. By the end of this you can read every box.

A 20-minute read 10 demos you can click No maths required
Before we start · who this is for

No maths, no jargon, and nothing you have to take on trust.

If you have only used chatbots

You will find out what is actually happening when one answers you, and why it behaves the way it does.

If you build with APIs

Tokens are the bill, the context window is the memory, temperature is the dice. All three are explained here.

If you want to go further

Every number on these slides is sourced, and the last slide has the commands to reproduce every demo yourself.

Every idea, no equations. Every demo below runs offline, on a laptop.
The hook · pop quiz

Pop quiz. Be honest with yourself, then read on.

Six mysteries. By the end of today, you can explain every one.

Pop quiz · pick one
You ask an AI chatbot to summarise a research paper that does not exist. It writes a confident, detailed summary anyway. Why?
The game · guess before you click

Beat the AI: guess the next word before it does

Prompt 1 of 2

Make your guess before you hit reveal. You are about to do, by hand, the only thing a language model ever does.

What just happened

The model's top 5 guesses · hidden
describes the animal describes the street

Real numbers: gemma3:4b, an open-weights model (4-bit, as Ollama ships it), recorded offline in raw next-token mode, September 2026. The top 5 of the 262,144 tokens it knows.

The plan · six parts

What is ahead: six parts, and you have already done the first

1
The hook
A pop quiz, six mysteries and a game of beat the AI
done
2
Where it came from
One paper in 2017, and nine years from translation to AI agents
3
Inside the machine half the deck
Tokens, meaning, attention, the block, the next word
4
Why it behaves like that
Forgetting long chats, and confidently making things up
5
Then and now
Same block, bigger models; every mystery solved
6
Run it yourself
Deep dives, plus the commands to reproduce every demo

Every idea, no equations. One sentence carried through all nine stations.

Four words you will hear all day

Open up a chatbot and you find four layers

In one line: an LLM is a next-word predictor built as a Transformer, and attention decides which earlier words matter.
The app you type into
Chatbot

ChatGPT, Gemini, Claude, DeepSeek: a chat window wrapped around an LLM that was also trained to follow instructions and hold a conversation.

ChatGPTGeminiClaudeDeepSeek
Open it up: what is inside the chatbot?
The model
LLM
Large Language Model

A program trained on a huge pile of text to do one thing: predict the next token. It never looks anything up. Large means billions of learned numbers; gemma3:4b, used in the demos here, has 4 billion.

Explains 236
What is inside the LLM?
The architecture, its design
Transformer
Vaswani et al., 2017

Text becomes tokens, tokens become lists of numbers, and the same block is stacked again and again. It reads chunks, not letters, and only what fits in its window.

Explains 15
What is inside each block?
The key idea
Attention

Every word looks at every other word, all at once, and picks the ones that matter. That is how it finds animal.

Explains 4
Where it came from

This is where everything started: one paper, in 2017

From a translation paper to AI agents in nine years

The first page of Attention Is All You Need, arXiv:1706.03762, showing the title and the eight authors
arXiv:1706.03762 · NeurIPS 2017 · real first page
over 190,000
citations · Semantic Scholar, 2026
12 h
to train the base model on 8 GPUs

Eight researchers wanted a better translator

They threw out the step-by-step reading of older models and kept only attention. That one idea now runs every chatbot you use.

The one-sentence pitch, from the abstract

"We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely."

First page reproduced from arXiv:1706.03762 (v7), Vaswani et al., 2017.
2017
The paper

A better translator, built only from attention.

Key open weights fully open: weights, data, code open source tool closed

Open weights is not open source. You can download and run these models, but the training data and code mostly stay hidden. OLMo publishes all three.

Open weights ≠ open source: Llama's licence has limits; Mistral 7B, DeepSeek-R1 and gpt-oss use Apache 2.0 or MIT; OLMo also publishes its training data.
Why learn this

Why spend 20 minutes on this? It runs every AI tool you touch, and you can open it up yourself.

Debug AI instead of guessing

Know why it hallucinates, forgets and miscounts, and work around it on purpose.

Build better, cheaper apps

Tokens are the bill, the context window is the memory, temperature is the dice.

Read any new model paper

Llama, DeepSeek, Qwen, Gemma: the same block with different settings. Try one tonight.

Run it, study it, change it

Open models run on a laptop. Free software needs people who understand what is inside.

Follow along

Every demo here runs on your own machine.

We use Ollama, a free tool that downloads an open model and runs it offline. Install it, pull one small model, and you can reproduce every demo in this deck. Where a demo can be re-run, the command is printed next to it.

ollama pull gemma3:1b   # 815 MB, runs on a laptop
Follow along · on your own machine

Every demo from here on, you can run yourself in two commands

Ollama is a free, open-source tool that downloads an open-weights model and runs it entirely offline. Nothing leaves your machine, and there is no account or API key.

1 · Install and pull one small model
ollama pull gemma3:1b   # 815 MB
ollama pull gemma3:4b   # 3.3 GB, optional

gemma3:1b is Google's smallest open model. It is deliberately small, which is exactly why it shows the rough edges so clearly.

2 · Ask it something
ollama run gemma3:1b "Why is the sky blue?"

That is the whole setup. From here, wherever a demo can be reproduced, the exact command sits on the slide next to it, marked run it yourself.

A fair warning. A 1-billion-parameter model on a laptop is far weaker than the chatbots you are used to. That is the point: the failures we are about to explain are easiest to see on a small model, and the mechanism behind them is identical in the large ones.

Before 2017, models read one word at a time, like passing a note down a line

RNN / LSTM: one word after another

0 steps
How much of animal is left when the note reaches it
·

Transformer: every word at once

0 steps
it links straight to animal: one hop, nothing lost.
Problem 1 · slow

Word 7 has to wait for word 6. A graphics card with 10,000 cores sits mostly idle.

Problem 2 · forgetful

Information from early words fades with every hop. Long sentences lose their beginning.

Table 1 from Attention Is All You Need: complexity, sequential operations and maximum path length
Hard mode · Table 1 of the paper

Maximum path length between any two words: O(1) for self-attention, O(n) for recurrence. That one column is the paper's whole argument.

The price: O(n²·d) work per layer. Double the text, quadruple the attention work.

Station 1: the model never sees letters. It sees chunks called tokens.

GPT-2's real token ids ↑
Try it: the real tokenizer of GPT-2, an open model from 2019, running in your browser
0 tokens · 0 characters · dashed = a raw byte, not a whole character
1Mystery 1

How many r's in strawberry?

strawberry

It gets three chunks and never the letters inside them, so counting the r's means guessing. gemma3:1b answered ten. Bigger models often get it right now, by guessing better, not by seeing letters.

run it yourselfollama run gemma3:1b "How many r's in strawberry?"
Why builders care

Tokens are what you pay for. API prices and context limits are all counted in tokens.

Hard mode · byte-pair encoding

Start from 256 raw bytes. Merge the most frequent neighbouring pair. Repeat about 50,000 times. GPT-2 ends with 50,257 tokens; Llama 3 uses 128,256 and Gemma 3 about 262,000.

Kannada "ಕನ್ನಡ ಭಾಷೆ" costs GPT-2 27 tokens for 10 characters.

Station 2: every token becomes a list of numbers, and similar meanings land close together

An illustrative map · a real space has hundreds of dimensions · click two words
Hard mode · the maths of "close"

cos(a, b) = a·b / (|a| |b|)

Real vectors: 768 numbers per token in GPT-2, 4,096 in Llama 3 8B. Position in 2017: add sine waves. In 2026: RoPE rotates Q and K by position.

How close? (cosine similarity)
Click two words on the map.
·
One catch: order

The numbers say what a word is, not where it sits. So every token also gets a position stamp.

Station 3: attention lets every word look at every other word and pick who matters

Which does it mean here: the animal, or the street?
In plain English

You just did attention yourself: you read the whole sentence, worked out which word it points to, and ignored the rest. A Transformer does this for every word, in every layer. Click any word to see its pattern.

The weights shown are illustrative, shaped like the heads in the paper's Figures 3 to 5. A trained model computes a set like this for every word at once.

Hard mode · what the arcs are

Each arc is one attention weight: how much it borrows from that word. The weights for one word always add up to 100%. These are illustrative, shaped like the real heads in the paper's Figures 3 to 5, and a real model computes a set like this for every word at once.

How does it find animal? Every word writes three sticky notes.

Key · from animal
"What I am about"
I'm a living thing. I can get tired.
Value · from animal
"What I hand over"
If you pick me, take my meaning: a four-legged animal.
Key · from street
"What I am about"
I'm a place. I can be wide or busy.
Query · from it
"What I'm looking for"
Who am I? Something that can be tired.
it vs animal: strong match it vs street: weak match animal's value flows into it
Like a library search

Your question (Query) is compared with every book's label (Key), and the best matches hand over their contents (Value). Not one book: a blend, weighted by how well each one matched.

Hard mode · three learned matrices

q = x·WQ   k = x·WK   v = x·WV

Why not compare the raw word vectors? "it" is not similar to "animal", it is relevant to it. WQ and WK learn relevance. Nobody sets them by hand: training does.

Attention in three steps: score, softmax, blend

Attention in four steps: score, scale, softmax, blend

Not just “it”: every word asks at the same time

Every word is a Query

We followed one word, it, through one row of this grid. The model fills every row at the same moment: each word asks its own question, and each word is also a Key and a Value for all the others.

12 words ask, 12 words answer: 144 scores at once

A 1,000-token chat means a million scores, in every head of every layer. Nothing waits for anything else, which is exactly the work a graphics card is built for.

Hard mode · the attention matrix

This grid is softmax(QKT/√dk): one row per query, one softmax per row, all rows in one matrix multiply. In a chat model (a decoder) the hatched half is masked: a word may only look back. So in GPT-style models "it" cannot see "tired"; the link is made later, when "tired" looks back at "animal".

In plain English

"it" asks a question, every word answers with a match score, softmax turns the scores into shares of 100%, and "it" takes that mix of their meanings. Those shares are the attention.

Hard mode · why divide by √dk?

Attention(Q,K,V) = softmax( QKT / √dk ) · V

A dot product over dk numbers grows like √dk. Without the division, big heads give huge scores, softmax bets everything on one word, and the gradient dies. Try it: turn scaling off and drag dk.

dk 64

Multi-head attention: each head looks for a different kind of link

Multi-Head
Attention
the orange box in Figure 1, on the cover

Illustrative roles, modelled on heads found in trained models (Clark et al., 2019). Many real heads are messier or redundant (Michel et al., 2019).

One head, one kind of link

Each head has its own Query, Key and Value, so each one learns to look for a different kind of link in the same sentence. Nobody assigns these jobs: they appear during training. These four are illustrative. Some real heads are this clean; many are messier.

Then the heads combine

All heads run at the same moment. Their answers are joined side by side, then mixed into one update for each word. The 2017 paper used 8 heads. Llama 3 8B has 32 in each of its 32 layers: 1,024 heads.

Hard mode · three multiplies, no loops

X is the sentence, one row per token. Each head has its own WQ, WK, WV: Q = XWQ, QKT is an n × n grid, softmax each row, times V. Then Concat(head1…head8)·WO mixes the heads. 2017 base: 512 numbers → 8 heads of 64. Llama 3 8B: 32 query heads share 8 K/V sets.

Attention lets words talk. The MLP lets each word think. Stack that pair again and again.

Attention: a group chat

Words share notes with each other. This is the only place they talk.

MLP: desk work

Each word goes back to its own desk and thinks about what it heard. About two thirds of the weights live here.

Residual stream: add, never erase

Each layer writes a small edit on top of what came before, so nothing is lost on the way up.

Stack it

6 blocks in 2017, 32 in Llama 3 8B, 61 in DeepSeek-V3. The same block, and the diagram on the cover.

Hard mode · inside one block

FFN(x) = W2 · max(0, W1x + b1) + b2

512 → 2048 → 512 in the base model. Each sublayer is wrapped as x + Sublayer(x), then LayerNorm ("Add & Norm").

Three wirings: encoder-only (BERT) · decoder-only (Llama, Qwen, DeepSeek, Gemma, GPT) · encoder-decoder (the 2017 original, T5, Whisper).

Figure 1 of Attention Is All You Need: the full Transformer architecture

Figure 1 of the paper. You can now name every box in it. Vaswani et al., 2017, reproduced with the paper's stated permission.

Station 5: the real job is to hide the last word and predict it

Temperature 1.0

gemma3:4b's real top 8 for this sentence, rescaled to add up to 100%. Rolls are seeded, so the demo repeats.

Everything so far was preparation

Every station turned each word into numbers full of context. The model's real job, the game from slide 4, is to guess the word that comes next.

So hide tired. The last word, too, now carries the whole sentence. With tired gone, "it" could be the animal or the street again, so the guesses split.

Rolls so far
Temperature is a randomness dial

Low: it always takes the top word, safe but repetitive. High: long shots win and nonsense creeps in. It is not a creativity dial.

Hard mode · training on every position

pi = ezi/T / Σj ezj/T

While writing, token n may only look back at tokens 1 to n: the causal mask. That is also why training can predict every position of a sentence at once.

A chat reply is typed one token at a time: predict, add it, run everything again

Chat · the reply arrives one token at a time
Inside the model, right now
1 It reads everything so far: your message plus every token it has written
2 It scores every token in its vocabulary, all 262,144 it could write next. The top 5:
3 The winner is added to the end of the input, and everything runs again
Why one at a time?

Each new word depends on every word before it, including the model's own last word. Word 5 cannot be guessed until word 4 exists.

Hard mode · the KV cache, measured

The past never changes, so its keys and values are kept. Each new pass only computes the newest token.

Station 6: the model only sees what fits in its context window

The same chat, with a tiny window so it overflows fast
Window size
What is a context window?

It is everything the model can read at once: your messages, its replies, all of it, counted in tokens. When the chat outgrows it, the oldest part falls out. No error, no warning.

Real context windows, in tokens

Each lab's model docs, 24 Sep 2026. 1M tokens ≈ 2,000 to 3,000 pages.

Hard mode · why windows are expensive

Attention compares every token with every token: O(n²). Double the window and that work quadruples, and the KV cache grows with every token.

"Lost in the middle" (Liu et al., 2023): facts buried mid-context get used less. A tendency, strongest in smaller and older models.

Station 7: why it makes things up. There is no lookup, only the most likely-sounding words.

2Remember the quiz?
The model did its job

A plausible abstract is the most likely continuation. Truth was never part of the objective. Bigger models with search and tools do this far less: same mechanism, lower rate.

run it yourselfollama run gemma3:1b "Summarise the 2019 paper 'Attention Fields in Sparse Halting Networks' in three sentences."
Hard mode · why nothing stops it

Training minimises next-token error. Nothing in that rewards "I don't know". OpenAI's 2025 paper Why Language Models Hallucinate argues training and benchmarks reward confident guessing. Fixes: retrieval, tools, training to abstain.

Prompt
Summarise the 2019 paper "Attention Fields in Sparse Halting Networks" in three sentences.
gemma3:1b · open weights · confident

"...introduces a novel method for approximating the outputs of sparse halting networks, which are used in reinforcement learning … achieving impressive results in tasks like Atari games."

gemma3:4b · open weights · just as confident

"The paper introduces Halting Networks, a novel architecture for predicting whether a program will halt, using a sparse, learned attention mechanism over the execution trace."

Real outputs, same prompt, recorded offline. Two models, two completely different inventions. If they were looking something up, they would agree.

DOES NOT EXISTno such paper, no such result

2017 to 2026: the block is almost unchanged. The scale is not.

Totals as each developer states them on its model card. Closed labs (GPT-4 onward, Claude, Gemini) no longer publish sizes, so they cannot be plotted.

The constant

Attention, then MLP, with a residual stream, stacked. The same block you just built.

The trick since 2024

Many small expert MLPs, and a router picks a few per token. Kimi K3 (2026, open weights) has 2.8 trillion weights but uses only 104 billion for each token.

Hard mode · five swaps since 2017

RoPE: rotate Q and K by position instead of adding sine waves.
KV cache: never recompute the past.
GQA: many query heads share one set of K and V.
FlashAttention: same maths, tiled so the n × n grid never hits memory.
Mixture of experts: many MLPs, a router picks a few.

Read any model paper with four questions: how do tokens talk, what is the MLP, how is position encoded, what is it trained to predict?

And the tokens do not have to be words

Images · 16×16 patches (ViT)

Audio · slices of sound (Whisper)

Proteins · amino acids (AlphaFold)

Code · a variable finds where it was declared

Recap · mysteries solved

Six mysteries, one idea: predict the next token, and let attention decide what matters

If you remember one thing: it predicts the next token, and attention decides what matters.
Where to go next

Now go and open one yourself.

Go deeper · all free
poloclub.github.io/transformer-explainera live GPT-2 in your browser: hover any token, watch attention light up
tiktokenizer.vercel.appthe exact tokens and ids for real models
bbycroft.net/llma 3-D walk through every matrix inside GPT
3Blue1Brown · "Attention in transformers"the visual QKV maths, done properly
arxiv.org/abs/1706.03762the paper itself. It will read differently now.
huggingface.co/modelsthousands of open-weights models to download and study
allenai.org/olmoa fully open model: weights, training data and code

Try it yourself

1. Install Ollama (free, MIT licence) and run ollama run gemma3:1b
2. Also try llama3.2, qwen3, deepseek-r1: all open weights
3. Ask one about a paper that does not exist
4. Then Karpathy's "Let's build GPT": 2 hours, about 300 lines

Deep dives

Softmax, worked out · how Q, K and V are learned · the full command sheet · thousandmiles.ai

First given as a live session for engineering students new to LLMs, then rebuilt to read on its own.

Deep dive · on request

Softmax, with the actual arithmetic

softmax(z)i = ezi / Σj ezj

Two properties, which is why it is used

Exponentiating makes every number positive. Dividing by the sum makes them add to 1. So any list of scores becomes a valid set of weights.

Why exponentiate at all?

It leans towards the winner. A gap of 2 in the scores becomes a factor of about 7.4 in the weights, so attention stays selective instead of a bland average.

Why it trains

It is smooth everywhere, so gradients flow back into WQ and WK. A hard "pick the max" would give no gradient at all. Real code also subtracts the max score first so ez never overflows.

Deep dive · on request

Nobody told it that "it" means "animal". So how did WQ and WK find out?

0

At the start it is noise. WQ, WK, WV are random, queries match keys at random, and the guesses are gibberish.

1

One thing is measured. Show it "The animal didn't cross the street because it was too", ask for the next word, compare with the real one. One number: how wrong.

2

The error flows backwards. From the loss, through the output, the blend, the softmax and the dot products, into WQ and WK. Every step is differentiable, which is why attention is built from dot products and a softmax.

3

A small nudge. "This guess would have been better if 'it' had looked more at 'animal'." So the matrices shift to line that query and key up a little. Repeat over billions of tokens.

→

Grammar falls out as a side effect. Pronoun resolution was never a goal. It is simply what helps predict the next word, and the same pressure produces heads nobody has a name for.

Run every demo yourself · command sheet

Everything here runs offline on a laptop

Setup
ollama --version        # needs 0.12.11+ for logprobs
ollama pull gemma3:1b   # ~800 MB
ollama pull gemma3:4b   # ~3.3 GB, the numbers in this deck
1 · Next-token probabilities (Beat the AI)
curl localhost:11434/api/generate -d '{
  "model": "gemma3:4b", "raw": true, "stream": false,
  "prompt": "The capital of France is",
  "logprobs": true, "top_logprobs": 5,
  "options": {"num_predict": 1}
}'   # probability = exp(logprob)
2 · Temperature
ollama run gemma3:1b
/set parameter temperature 0     # ask 3 times: identical
/set parameter temperature 1.5   # ask 3 times: all different
3 · Context window forgetting
/set parameter num_ctx 256
# "The secret code is BANANA." + 300 tokens of filler
# then ask for the code: it cannot say
/set parameter num_ctx 4096      # now it can
4 · Hallucination
ollama run gemma3:1b "Summarise the 2019 paper
 \"Attention Fields in Sparse Halting Networks\"
 in three sentences."
Two gotchas

Pre-warm with "keep_alive": -1 or the first call stalls while the model loads. Use "raw": true for pure next-token numbers, or the chat template changes them.

thousandmiles.ai
EASY HARD
01 / 20
1 / 20