JournalNLP
Tokenization Decides What Your Model Can Count
The vocabulary you never think about sets a hard ceiling on arithmetic, multilingual cost, and how well Urdu is treated.
Tokenization looks like plumbing. You pick a byte-pair encoding, train it on a corpus, and move on to the parts that feel like real modelling. In practice the vocabulary you choose quietly sets a ceiling on everything downstream — and unlike most ceilings, no amount of extra parameters raises it.
The arithmetic tell
The clearest symptom is numbers. When 1,024 splits into three unrelated pieces, the model must learn place value across boundaries that carry no positional meaning. It is being asked to do arithmetic through a keyhole.
"1024" → ["10", "24"] # BPE, merged by frequency
"1024" → ["102", "4"] # a different corpus, different merges
"1024" → ["1","0","2","4"] # digit-level: place value is learnableDigit-level tokenization fixes more arithmetic failures than another few billion parameters usually will. It costs you sequence length and buys you a representation where the units column is always the last token.
The multilingual tax
Outside English the effect sharpens. Morphologically rich and non-Latin scripts routinely burn two to four times as many tokens per word. That is not an abstract inefficiency — it compounds in three directions at once:
- A shorter effective context window, because the same document consumes more of it.
- Higher inference cost per user, for identical content.
- Worse sample efficiency, because meaning is smeared across more positions during training.
For Urdu specifically, a vocabulary trained mostly on English fragments words in ways that destroy the morpheme boundaries a reader depends on. The model is not reading Urdu badly; it is reading a different language that happens to share the bytes.
Measure tokens-per-word on the languages you actually serve before you tune anything else. It is a two-line diagnostic that explains results people spend weeks blaming on model quality.
What to do on Monday
- Compute the tokens-per-word ratio for each language in your traffic mix.
- If any language exceeds roughly 1.5× your English baseline, price and budget context for it separately.
- If numeric reasoning matters, test a digit-level variant before touching the architecture.
- Where you control training, extend the vocabulary with the target script rather than relying on byte fallback.