Why Non-English Costs More: What a Tokenizer Actually Decides
Published October 9, 2026
One tokenizer can turn the same word into 5 tokens or 15. That difference decides your API bill and how much context you actually have.

One word, two token counts
Take the same sentence written twice, in two languages. The English version might take five tokens. The translated version can take fifteen. The meaning is the same, but a model does not read sentences. It counts pieces called tokens.
A tokenizer is the component that turns text into those pieces. It does not read meaning, and it does not know which words are common. It has a dictionary of pieces, and each piece has a number.
That is why API billing is not counted in characters or in words. APIs count tokens. And the token count for the same content can differ a lot between languages.
Why tokenizers are built this way
Most modern tokenizers are built the same way: the BPE algorithm, or Byte Pair Encoding, trained on a large korpus. The process repeats: take the most frequent piece, make it a new unit, repeat until the vocabulary size is reached.
The result resembles a dictionary of highly specific pieces. The word "the" becomes one token. "nonetheless" becomes two. A combination that rarely appears in the training korpus gets split into several pieces that are each already known.
The training korpus decides the result. If that korpus is dominated by English, the tokenizer will be very efficient for English and only moderately efficient for other languages. That is exactly the problem.
A calculation you can run yourself
The clearest way to understand this is to count. Take one paragraph in English, count its tokens, translate it to Indonesian without changing its length, then count again.
| Paragraph | Language | Expected tokens | Note |
|---|---|---|---|
| One short sentence | English | low | Common words become one token |
| One short sentence | Indonesian | higher | Words split more often |
| One technical paragraph | English | medium | Long terms become several tokens |
| One technical paragraph | Indonesian | high | Same terms, more splitting |
| Numbers and code | both | high | Numbers often split per digit |
The pattern is consistent: the less common a word pattern is in the training korpus, the more pieces get used. For languages composed differently from English, this is not a rare edge case.
Numbers that need their status attached
Many efficiency claims about tokenizers come from vendor testing. The metric used is often compression ratio, meaning characters per token. That number is useful, but only for one language and one domain.
What to ask is whether the test included non-English languages, and whether the measurement was done on the same text or on a translation. Testing a translation produces a different number from testing the original, because sentence lengths differ.
Analysis: where the real cost appears
Tokenizer cost shows up in three places, and all three can be counted.
First, API cost. If one request in Indonesian uses 1.6 times the tokens of the same task in English, the cost per request rises accordingly. This is not overhead; it is the rate.
Second, context length. Models have a context limit counted in tokens. If your technical documentation in Indonesian uses 1.6 times the tokens, only about 62 percent of the English version fits. Reading twice as many tokens means less fits at once.
Third, prompt caching. Many APIs cache repeated prefixes. If that prefix is in a token-hungry language, the reusable part shrinks and the effective cost rises.
| Where it shows up | English | Indonesian | Rough ratio |
|---|---|---|---|
| Tokens per paragraph | baseline | higher | 1.4 to 1.8 |
| Documents fitting in context | more | fewer | about 0.6 |
| Cost per request | baseline | higher | matches token ratio |
| Usable prefix cache | larger | smaller | lower |
What you can do
First, count before claiming. Most language cost problems are solved by counting tokens per language on the text you actually send. Without numbers, there is nothing to claim.
Second, compress the prompt, not the language. Removing unnecessary words helps in every language, because it cuts the same repeated pieces.
Third, watch the format. Numbers, IDs and code are often more token-hungry than ordinary words. If your requests contain long numbers, check the format before blaming the language.
Additional context: why character count is a bad proxy
Character count looks like a fair measure and is not. The same character count can produce very different token counts depending on which characters it is.
| Content | Characters | Tokens | Why |
|---|---|---|---|
| English prose | 1000 | about 220 | Common words, whole tokens |
| Indonesian prose | 1000 | about 340 | More splitting |
| Base64 payload | 1000 | about 700 | Rare pieces, heavy splitting |
| Compressed text | 1000 | about 900 | Almost no repeated pieces |
| Repeated prompt | 1000 | about 200 | Cache-friendly if prefixed |
The last two rows explain why a request that "looks the same size" can cost several times more. Compressed or encoded data is close to worst case, because the tokenizer has nothing to reuse.
This is also why shortening a prompt by characters rarely helps much. Cutting from 1000 to 700 characters may remove very few tokens if the removed parts were the well-tokenized prose. Cutting repeated filler removes whole pieces instead, and that is where the saving is.
| Method | Typical saving | Reliable |
|---|---|---|
| Cut characters blindly | low | No |
| Remove filler words | medium | Yes |
| Merge repeated sentences | medium | Yes |
| Move fixed text to the prefix | high | Yes |
| Switch to a compact format | high | Yes |
Move fixed instructions to the front of the prompt and the cache can reuse them. That single change often matters more than any wording tweak, and it applies to every language equally.
Additional context: why a similar language is cheaper
This is about the language, not the country. Indonesian uses reduplication, as in "anak-anak" or "buku-buku". That pattern is very rare in an English corpus, so a tokenizer has to split it into several pieces instead of one.
This is not an edge case. Malay, Javanese and most Austronesian languages share the same repetition pattern. English has a simple plural with an added "s" and a single word for negation. Those two languages produce very different pieces for the same concept.
The practical result is a larger number of distinct word forms in Indonesian for the same set of meanings. The tokenizer does more work and produces more pieces.
One simple measurement needs no special tool and no estimating. Take a paragraph you have already written, count its tokens, then count its characters. Multiply the character count by one hundred and divide by the token count. What comes out is your average characters per token.
For English that figure usually sits around four to five characters per token. For Indonesian it commonly lands between three and four. That gap is not a curiosity. If your prompt runs to 100,000 characters, English spends roughly 21,000 tokens while the Indonesian version spends roughly 28,000. Those seven thousand tokens are enough to cut an entire attached documentation section.
The second measurement surprises people more often. Take the same text, count tokens in both languages, then divide the Indonesian token count by the English one. For narrative text the ratio usually falls between 1.4 and 1.8. For text packed with English technical terms it drops toward 1.1. For numeric data and code it approaches 1.0.
The implication is that the language gap is not a constant. A technical document can cost almost the same in either language, while an explanatory document can cost far more in the differently composed language. Knowing your own ratio is more useful than adopting somebody else's constant.
| Phenomenon | English | Indonesian | Effect on tokens |
|---|---|---|---|
| Reduplication | absent | common | more pieces per word |
| Plural | one letter "s" | more complex | affix becomes its own piece |
| Negation | "not" | "tidak", "bukan", "nggak" | several distinct pieces |
| Question reduplication | absent | common | one meaning, several forms |
| Tense marking | simpler | many forms | more unique forms |
A short history: from words to pieces
Early tokenizers were simple: one word became one token, with a dictionary of a few tens of thousands of entries. The problem showed up immediately. A word not in the dictionary got split per character, so one long word could cost ten tokens or more.
BPE arrived as the answer. Instead of storing whole words, BPE stores pieces. "Baik" and "mau" can share the ending. An unknown word no longer breaks into letters, but into pieces that are already known. In practice this lowers the token count on unexpected text.
Work after BPE focused on tokenizers designed for languages other than English. That work is still early. Most large models still use tokenizers trained on English-dominant corpora, and the efficiency gap between languages has not closed.
What remains is not purely a technical problem. Retraining a tokenizer means retraining the model, or at least reworking every embedding around it. That is expensive, and the payoff only becomes clear when enough non-English users exist to notice it.
Additional context: a real per-language ratio
The numbers above can be checked rather than estimated. Take 200 English words, count the tokens, translate, and count again. The gap usually lands between 1.4 and 1.8 times for narrative text.
Status note: unverified. The characters-per-token figures and the 1.4 to 1.8 ratio in this article are estimates, not measurements on a specific tokenizer. They change with the model's tokenizer, the text domain, and how the translation was produced. The reliable method is to measure your own text, since several APIs return the token count in every response.
One exception is often missed: text full of proper names, abbreviations and untranslated technical terms. That mix actually splits less in both languages, because proper names are often already single tokens in the vocabulary. So the language gap is not always as large as people assume.
Three things raise the ratio, in order of how often they appear:
| Cause | Effect on tokens | How often it appears |
|---|---|---|
| Reduplication | 2 to 4 extra tokens per word | Very often |
| Repeated affixes | 1 extra token per affix | Often |
| Numbers, dates, code | Little change | Always present |
| Untranslated proper names | Almost no change | Often |
The first two rows account for most of the gap. The last two explain why some documents cost nearly the same in either language.
A worked example: one document, two bills
Take a technical document of 5,000 characters. The English version and the translated version carry exactly the same information.
| Document content | Characters | English tokens | Indonesian tokens | Difference |
|---|---|---|---|---|
| Conceptual prose | 5000 | 1100 | 1700 | 55 percent more |
| API list with JSON | 5000 | 2100 | 2400 | 14 percent more |
| Numbers and code table | 5000 | 2400 | 2500 | 4 percent more |
| Compressed text | 5000 | 3600 | 3600 | the same |
The pattern in that table matters: the gap is not uniform. The language is not always the most expensive one. Prose is the expensive case, because that is where the repeated word patterns are lost. Numbers and code cost nearly the same in both languages, because there are few language patterns to reuse.
Now put that 5,000 character document into an 8,000 token context alongside the conversation history. The English version still fits. The Indonesian version, at 1,700 tokens, leaves far less room. In a system with long instructions, that room decides whether the instructions are still read whole or already truncated.
| System | Instructions | Document | English total | Indonesian total | Fits |
|---|---|---|---|---|---|
| A | 500 | 5000 characters | 1600 | 2200 | Yes |
| B | 1500 | 5000 characters | 2600 | 3200 | Yes |
| C | 3000 | 5000 characters | 4100 | 4700 | No, over 8k |
System C is the common case. The instructions are already long, the attached document is larger, and together they pass the limit. What usually gets dropped is the document rather than the instructions, because a document can be trimmed without breaking the answer, while a truncated instruction changes what the model was asked to do.
Three steps before changing models
Before changing models, three cheaper steps come first.
First, count tokens per language on your own text. Take ten paragraphs from the work you actually send, count them in English and after translation. That number beats any tokenizer benchmark.
Second, move fixed text to the front of the prompt. If instructions or examples are always the same and sit at the start, providers supporting prompt cache can reuse them. This change is language-independent and often larger than any prompt trimming result.
Third, tidy numbers and code. Use one consistent date format, avoid long unbroken numeric strings, and do not wrap data in base64 unless it is genuinely required. All of it removes pieces without removing information.
| Step | Difficulty | Expected saving | Needs a new model |
|---|---|---|---|
| Count tokens per language | Low | Indirect, but avoids a wrong decision | No |
| Move fixed text to the front | Low | Large when the prefix is long | No |
| Tidy numbers and code | Low | Moderate | No |
| Request a language-specific tokenizer | High | Large, but needs a new model | Yes |
The last row matters. Changing the tokenizer means changing the model, which is a far bigger decision than most people picture when they first learn that tokens exist.
Related tools
Free browser tools that apply to this topic.
- AI Writing TellsFind generic patterns in your own writing. Not an AI detector.
- AI Text HumanizerRewrite text into more natural, readable language while preserving its meaning.
- AI Text SummarizerCreate an abstractive summary locally in your browser with AI.
- Text and Line ToolsCount, dedupe, sort, filter, reformat, and convert case across twelve line operations in one box.
Share this article
Share to
Related articles

October 8, 2026
One Model or Many: What Multimodal Architecture Actually Changes
Unified models run text, image and audio through one network. Pipelines chain separate models. The difference is not only technical.

October 8, 2026
Why AI Benchmark Scores Are So Often Misleading
The numbers carrying a vendor's biggest claims are often measured on tests users never run. Four reasons, and how to read them.

October 8, 2026
Why Images Use Diffusion and Text Does Not
Two ways to make an image out of noise share the same endpoint and opposite processes. Why the choice has flipped since 2015.



