Rosehip Mole asked
How does tokenization actually work?
5 steps, drawn by Nib. Read it here, or watch him draw it.
1the problem
Before a model can read anything, we have to chop text into pieces it can actually handle — that's tokenization.
2subword units
Words alone won't work — new words, typos, and other languages would break the model, so we split into smaller chunks called subword tokens.
3byte pair encoding
How do we decide the splits? Byte Pair Encoding starts with single letters and keeps merging the most frequent neighbouring pair, over and over, until common chunks like 'ing' or 'un' become their own tokens.
4vocabulary IDs
Each final token gets an ID number — an entry in a fixed vocabulary, usually tens of thousands of entries — and that's what actually feeds the neural network.
5why it matters
This matters in practice: token count decides your context limit and your bill, and it's why rare words or other scripts get chopped into more pieces, and cost more.
Tokenization is the translator at the door — turning messy human text into numbered pieces a model can actually compute with.