Tokenization: A Survey for Modern NLP

· arXiv ·

3 min read Original article ↗

Language models process token IDs, not raw text. Converting text into those units has consequences for what models process and how much computation they use. The choice of units can also change how much processing equivalent text requires across languages. A recent survey maps how tokenizers are built, what can go wrong, and how to measure their effects in modern language models.

Flow diagram from normalized text through subword tokens and integer IDs to embedding vectors
The tokenization pipeline: normalization, pretokenization, tokenization, ID mapping, and embedding. The example splits "tokenization" and "pipeline" into pieces. This is a schematic example, not a quantitative result.

The text-to-model interface

A tokenizer is the vocabulary plus the procedure that maps text to tokens. The process runs in stages:

  1. Text may be normalized and divided into coarse pre-tokens.
  2. Tokenization splits it into vocabulary pieces.
  3. Those pieces map to integer IDs and then to embedding vectors.

Subword tokens are units between characters or bytes and whole words. They can represent unseen word forms without a word-level vocabulary large enough to cover every form. They also produce shorter sequences than character-level input.

The choice involves a trade-off. A vocabulary that is too broad costs parameters in the embedding and output layers. More granular tokens mean longer sequences, which can increase computation. Table 2.1 gives one illustration: the tied-embedding Gemma 4 E2B has a 262k-token vocabulary and allocates 59.80% of its 4.65B total parameters to the embedding and output layers.

Scatterplot of corpus token count against pretraining data fraction, with one point per language

Frequency-driven vocabulary building tends to favor languages better represented in the tokenizer's training data. A shared vocabulary may therefore encode some languages compactly and split others into more pieces. The survey calls this oversegmentation: breaking text into more, often less meaningful, subword units than a better-supported language receives.

More tokens can mean longer model sequences and higher latency or resource use. Standard transformer self-attention's computational cost grows quadratically with sequence length, but CTC itself is only a proxy for this burden.

Measuring these effects is difficult. Next-token cross-entropy and perplexity are token-level quantities. They depend on how many tokens encode a given string, so they cannot be naively compared across tokenizers. Bits-per-byte normalizes by byte count instead. Byte counts also vary with writing system, however, so bits-per-byte is not automatically a fair cross-language measure.

Tokenizer ablations can require expensive, noisy from-scratch model training. The survey reports that, in 2025, papers mentioning tokenization-related terms made up at most 2% of papers in the venues counted.

The authors do not propose a single replacement tokenizer. Their response is to:

  • improve measurement and benchmarks,
  • test effects at relevant scales, and
  • connect intrinsic properties such as compression to downstream performance.

Fewer tokens can reduce token-metered charges, all else equal. Compression does not by itself prove better model quality, faster wall-clock inference, or lower overall cost.

The survey's contribution is to connect tokenization to practical effects: compute and inference costs, parameter budgets, and unequal treatment of languages. It also calls for better evaluation so tokenizer choices can be studied more confidently. The authors present this as a case for more research, not as evidence that one proposed tokenizer is best.