Back to News & insightsGuides

Multilingual AI: the hidden product effects of tokenization

Why identical character limits can create unequal experiences across languages, and how to test budgets, truncation, and quality fairly.

Editorial guide · Updated September 27, 2026 · 4 min read
Different segmented ceramic shapes travel through parallel silver channels.

A text box that accepts the same number of characters from everyone looks fair. Behind it, a model may turn those characters into very different numbers of tokens. The difference can affect how much material fits, how often users hit a limit, and how expensive the complete interaction becomes.

For an international AI product, tokenization belongs in product planning as well as infrastructure planning. English examples alone cannot establish a suitable limit for every language your interface supports.

Words, characters, and tokens are different units

A tokenizer maps text into units used by a model. Those units do not have a universal relationship to words or characters. Research on language-model tokenizers has documented substantial differences in tokenization across languages for the systems studied.

Treat that research as a reason to measure your chosen tokenizer, not as a permanent ranking of languages or models. Tokenizer revisions, writing systems, punctuation, code, and the kind of content being processed all affect the workload.

Start with equivalent tasks

Imagine a hypothetical museum assistant that explains exhibits in several languages. Comparing one short English sentence with a long visitor essay in another language would reveal little. Build task-equivalent examples: the same exhibit description, the same visitor question, and answers of comparable usefulness.

Have fluent reviewers check those examples. Literal translations can change tone, ambiguity, and the amount of explanation needed. The goal is equivalent meaning and purpose, not an artificial requirement that every response have the same number of words.

Keep naturally written examples too. Visitors may mix languages, use local names, or paste text with unusual spacing. A translated test set is a controlled experiment, while native submissions help reveal the actual product experience.

Measure the complete request

Count tokens for the system instructions, supporting evidence, conversation history, and expected output. Measuring only the visible user message can understate the budget. Record the exact tokenizer or provider usage report used for the calculation.

For the museum assistant, compare a first question with the fifth turn of the same conversation. History may consume more space than the latest question. Also test proper names and exhibit identifiers, which can behave differently from ordinary prose.

Track distributions rather than a single average. A small group of unusually expensive requests can determine when the interface fails. Keep quality results beside usage results so a configuration is not selected simply because it produces shorter, less helpful answers.

Design limits around the task

Suppose a visitor reaches the context limit halfway through a discussion. Silently deleting the exhibit evidence can leave the model sounding confident while losing the facts it needed. A better policy identifies which context is essential and makes any reduction explicit.

Possible approaches include shorter retrieved passages, a clearly marked conversation summary, or asking the visitor to start a new topic. Each needs evaluation. A summary that drops a negation or mistranslates an object name can change the answer even while reducing token use.

Show limits in terms users can understand, such as the amount of document content the service can accept. Internally enforce the actual token bound. Avoid promising that a fixed page count always fits when pages differ dramatically in language, layout, and density.

Evaluate meaning separately from fluency

A fluent localized answer may still mishandle a date, transliterate a name inconsistently, or answer in the wrong language after several turns. Ask reviewers to check factual preservation, instruction following, and appropriate wording separately.

For retrieval, test whether a question in one language finds evidence written in another. Otherwise, poor search can be mistaken for poor generation. Include a case where the correct evidence is supplied directly; the difference helps isolate the failing stage.

Make error reporting possible in the visitor's language. Feedback that only accepts an English explanation will miss some of the people most affected by the problem. Preserve enough context to investigate while avoiding unnecessary personal information.

Maintain a language-specific operating picture

Record rejection rates, response length, usage, latency, and reviewed quality by supported language where collection is appropriate. Investigate gaps rather than assuming all differences are caused by tokenization.

The practical objective is equal access to a useful task, not identical token consumption. A sound international product makes its tradeoffs visible and verifies that each supported language remains usable within the service's limits.

Research background

Read the original research paper on arXiv

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.