Machine learning

Why can’t AI count the r’s in “strawberry”?

Because it never sees the letters. It sees tokens.

the short answer

Language models don’t read letters. Before a model sees your question, a tokenizer chops the text into chunks called tokens and replaces each with a number: in GPT-4o’s tokenizer, the word “strawberry” inside a sentence, together with the space before it, is the single number 101830. The letters inside that number are invisible to the model; it can only know them by having learned them, so counting them is easy to get wrong.

Type “How many r’s are in strawberry?” into a chatbot, and for a long while the famous answer was two.2 The models could write sonnets about strawberries and explain how they grow, and still miscounted three letters. The reason is that the model never saw the letters at all.

Ask it as · read by

Before the model reads anything, a tokenizer chops your text into tokens, common chunks of characters, and swaps each for its number in a fixed list. To GPT-4o, “How many r's are in strawberry?” is 31 characters but just 8 tokens, and “ strawberry”, with the space in front of it, is the single token number 101830. That number is all the model gets. Whatever letters are inside it, the model has to have learned, the way you might learn the spelling of a word you’ve only ever heard.

Here’s a puzzle. On its own, at the very start of a message, “strawberry” has no space in front. How many tokens is it then?

Three, in all three tokenizers. GPT-4 cuts it as str · aw · berry, and GPT-2 and GPT-4o as st · raw · berry. Same word, different chunks, depending on a space. Try the examples above: capitals change the chunks again.

Where do tokens come from?

Nobody writes the list of tokens by hand. It is learned from text, by a trick first invented for compressing files called byte pair encoding.3 Start with single characters. Find the pair of neighbours that appears most often in a big pile of text, say t followed by h, and glue it into one new token, th. Then do it again, and again. After enough merges, common words are single tokens, and rare ones are built from pieces.4 Below, a tokenizer is learning from the first three chapters of Alice in Wonderland, right now, in your browser.

With 200 merges, your sentence is 16 tokens for 28 characters. Drag the number down to zero and the tokenizer sees letters, like you do. Drag it up and words it has read a lot (Alice, the, said) become single tokens, while words it has rarely seen stay in pieces.

Then why not just give the model letters?

Because it would cost far more. In English one token is about four characters,5 so reading letter by letter would make every text about four times longer, and a model’s work grows faster than the length of what it reads. Tokens let a model take in a whole book. Real tokenizers have a lot of tokens to choose from: GPT-2 has 50,257, GPT-4 about 100,000, GPT-4o about 200,000.

So how do models spell at all? From text that spells words out, they learn which letters each token holds, a memory rather than a look. Ask a model to write “strawberry” one letter at a time first, and it usually counts correctly: written out, each letter becomes its own token, and now the model can see them.

Wait, but then…

In GPT-2’s tokenizer, “ SolidGoldMagikarp”, a Reddit username, is one single token. In 2023 two researchers found that GPT-3, whose tokens are the same, went haywire when asked to repeat it: it evaded, insulted them, or said a different word.6 The likely reason: the token was in the list, learned from one pile of text, but hardly ever appeared in the text the model itself trained on. GPT-4’s tokenizer splits it into five ordinary pieces. What do you think a model has learned about a token it has almost never seen?

Try it yourself

Each task ticks itself off when the page above shows it.

  1. Find a way of writing “strawberry” that all three tokenizers cut into four pieces.In capitals: GPT-4 and GPT-4o cut it ST · RAW · B · ERRY, and GPT-2 ST · RAW · BER · RY. Capitals are rarer in the text a tokenizer learns from, so the shouted word falls into more, smaller pieces.
  2. Make the Alice tokenizer see letters, the way you do.With no merges, every character is its own token: easy to count, but four times as long.
  3. Type a single word of your own in the box, and add merges until it becomes one token.It happens only for words the tokenizer saw often in its training text; others stay as pieces however many merges you allow.
  4. In your own words: why can a model write about strawberries but miscount their r’s?

So AI miscounts the r’s in “strawberry” because the word reaches it as one or a few numbered chunks, not ten letters: it can only count what it remembers is inside them.

Sources

  1. OpenAI, tiktoken (open-source tokenizers): every real tokenization on this page was computed with tiktoken 0.9.0 using gpt2 (GPT-2), cl100k_base (GPT-4) and o200k_base (GPT-4o).
  2. Amanda Silberling, “Why AI can’t spell ‘strawberry’”, TechCrunch, 27 August 2024.
  3. Philip Gage, “A New Algorithm for Data Compression”, The C Users Journal 12(2), 1994: byte pair encoding.
  4. Rico Sennrich, Barry Haddow and Alexandra Birch, “Neural Machine Translation of Rare Words with Subword Units”, ACL 2016: byte pair encoding for language models.
  5. OpenAI Help Center, “What are tokens and how to count them?”: for English, one token is about four characters.
  6. Jessica Rumbelow and Matthew Watkins, “SolidGoldMagikarp (plus, prompt generation)”, LessWrong, February 2023.
  7. Lewis Carroll, Alice’s Adventures in Wonderland (1865), chapters I–III, from Project Gutenberg eBook #11 (public domain): the text this page’s own tokenizer learns from.

Written and drawn by Nib. Every drawing is drawn live and redrawn as you change it.