Dutch needed 20% more tokens than English to say the same thing. A token is the chunk of text an AI model reads, and the unit you pay for: a short word, or a piece of a longer one. We counted all 78 English and Dutch article pairs published on this site with the tokenizer (the program that cuts text into tokens) that OpenAI’s own counting library assigns to GPT-4o through GPT-5. English came to 90,765 tokens, Dutch to 109,235. AI services charge per token, so the same work in Dutch costs about a fifth more. On the older tokenizer that OpenAI names for its text-embedding-3 models, the gap was 49%.
What exactly did we measure?
On 24 September 2026 we took every article on lenvanderhof.com that is published in both English and Dutch: 78 pairs. Each Dutch version is a native rewrite of the same content by the same author, not a word-for-word translation.
We kept only the article body. We removed the metadata header and the web addresses inside links, but kept the link text. Then we counted with tiktoken 0.14.0, OpenAI’s open-source token counter, using two encodings (an encoding is a tokenizer’s fixed list of pieces):
- o200k_base, which tiktoken assigns to GPT-4o, GPT-4.1, GPT-5 and the o1, o3 and o4-mini models.
- cl100k_base, which tiktoken assigns to GPT-4 and GPT-3.5-turbo, and which OpenAI’s embeddings guide names for its text-embedding-3 models.
| All 78 pairs | English | Dutch | Dutch ÷ English |
|---|---|---|---|
| Words | 70,543 | 72,071 | 1.02 |
| Characters | 408,399 | 456,723 | 1.12 |
| Tokens, o200k_base | 90,765 | 109,235 | 1.20 |
| Tokens, cl100k_base | 91,259 | 135,540 | 1.49 |
| Tokens per word, o200k_base | 1.29 | 1.52 | |
| Tokens per word, cl100k_base | 1.29 | 1.88 |
The Dutch articles were only 2% longer in words. Almost the whole gap comes from how the words get cut up.
It was not a few outliers. All 78 pairs cost more in Dutch. On o200k_base the ratio per pair ran from 1.11 to 1.39, with a median (the middle value) of 1.19. On cl100k_base it ran from 1.28 to 1.74, median 1.47.
Why does Dutch split into more pieces?
A tokenizer learns its pieces from a large pile of training text. Pieces that appear often get their own token. Everything else is built from smaller parts. If the pile holds far more English than Dutch, English words are more likely to be whole pieces.
Aleksandar Petrov and colleagues at the University of Oxford measured this across 200 languages, in a study for the NeurIPS 2023 research conference. The same text translated into different languages can differ in token length “up to 15 times”, they found, and tokenizers “are heavily influenced by the biases of the corpus source”, meaning the text they learned from.
Dutch adds its own twist: compounds. Dutch glues words together where English keeps them apart. Here is what the two encodings did with a few words in our own test:
| Word | o200k_base | cl100k_base |
|---|---|---|
| hospital | 1 token | 1 token |
| ziekenhuis (hospital) | 1 token | 4 tokens |
| disability insurance | 2 tokens | 2 tokens |
| arbeidsongeschiktheidsverzekering (disability insurance) | 7 tokens | 10 tokens |
Each word was counted with a leading space, the way it appears mid-sentence.
The newer encoding learned “ziekenhuis” as one piece. The long compound still breaks into seven.
OpenAI pointed in the same direction when it launched GPT-4o in 2024. Its new tokenizer used fewer tokens for all 20 example languages it showed; a German sentence dropped from 34 to 29 tokens, an English one from 27 to 24. Dutch was not on that list. Our numbers show the same move: the Dutch gap shrank from 49% on the old encoding to 20% on the new one.
Is the 20% just our writing style?
Our Dutch articles are rewrites, and one author wrote all of them. So we cross-checked on a public benchmark.
FLORES-200 is a set of 3,001 sentences from 842 web articles, translated by humans into 200 languages. It was built by Meta’s NLLB research team, and Petrov’s study used it too. We took the 2,009 sentences in its public dev and devtest sets, in English and Dutch, and counted them the same way.
| FLORES-200, 2,009 sentences | English | Dutch | Dutch ÷ English |
|---|---|---|---|
| Tokens, o200k_base | 52,409 | 65,174 | 1.24 |
| Tokens, cl100k_base | 52,983 | 84,599 | 1.60 |
On cl100k_base we got 1.60. Petrov’s published figure for Dutch on the same encoding is 1.59, so the counting method holds up. On the newer o200k_base, literal translations came out 24% dearer in Dutch, a little worse than our rewrites. If anything, our 20% headline is on the low side.
What does 20% mean in euros?
Hypothetical arithmetic for a Dutch business that runs its AI work in Dutch:
| If the same work in English would cost | In Dutch, at +20% | Extra per year |
|---|---|---|
| €500 a month | €600 a month | €1,200 |
| €5,000 a month | €6,000 a month | €12,000 |
| €50,000 a month | €60,000 a month | €120,000 |
This assumes you pay per token and your model’s tokenizer behaves like o200k_base. If the model also answers in Dutch, the output grows too, and output is usually the most expensive line on the bill. The price per token itself differs widely between models; Undominated.ai is the index I built to compare those prices.
There is a second cost that is not money. Every model has a context window: the maximum number of tokens it can read at once. At our averages, 100,000 tokens hold about 77,700 words of English but only about 66,000 words of Dutch. The same window fits less Dutch.
What about RAG and embeddings?
RAG (retrieval-augmented generation) is the setup where an AI looks up passages in your own documents before it answers. To find the right passages, the documents are first turned into embeddings: lists of numbers that let a search system find text with a similar meaning.
OpenAI’s embeddings guide says to count tokens for its text-embedding-3 models with cl100k_base, and that “Usage is priced per input token.” On that encoding our Dutch text needed 49% more tokens. So embedding a Dutch document library takes about one and a half times the tokens of the same library in English. If you cut documents into chunks of a fixed number of tokens, you also get about 49% more chunks to store and search.
The same guide lists a maximum input of 8,192 tokens per embedding. At our averages, that is about 6,300 words of English or about 4,400 words of Dutch. What is RAG? explains the rest of the setup.
What this measurement does not show
- Rewrites, not translations. Our Dutch articles are native rewrites. The FLORES-200 check suggests literal translation gives a slightly larger gap, not a smaller one.
- One tokenizer family. We measured OpenAI’s encodings through tiktoken. Other vendors cut text differently. Anthropic’s pricing page says its newer Claude models use a tokenizer that “produces approximately 30% more tokens for the same text” than its previous one. Tokenizers vary even inside one company, so measure the model you actually use.
- A site, not a benchmark. These are 78 article pairs about AI, search and founders, by one author. A legal archive or a chat log could land elsewhere.
- Newer models. tiktoken 0.14.0 maps model names up to GPT-5. It does not tell you which encoding a newer model uses.
- Text only. We counted article text. Real API calls add tokens for message formatting, tools and files.
Where does ROUTE fit?
Language is a cost setting, so manage it like one. That is where ROUTE helps. ROUTE is the framework from my book The Model Portfolio, which is available now. It is for running several AI models like an investment portfolio: each task goes to a named model for a named reason, and its cost is measured. The five steps:
- R · Register. List every model you use with its purpose, cost class, privacy class and retirement status.
- O · Objective typing. Write down what each task needs (depth of reasoning, format, speed, quality, privacy) before you pick a model.
- U · Utilize policy. Write the rules that send each request to a model, with backups.
- T · Track. Measure cost, speed and quality per route, meaning per type of task, so waste shows up before the invoice.
- E · Evolve. Promote, demote and retire models on a schedule.
You do not need the book to use this. Applied to Dutch, two steps do the work.
Objective typing: add language to the task description. A Dutch support route and an English one have different token budgets, even with the same model and the same prompt. The book also recommends tagging each model with what it handles well and badly, and calls a documented failure on non-English input “essential for any team routing multilingual traffic.”
Track: log tokens per task, per language. If Dutch tasks cost about a fifth more, that matches this measurement. If they cost far more than your own measured ratio, find out why before you scale.
Try this today (15 minutes)
Measure your own gap. You need Python and five documents you have in both English and Dutch. This is the method we used, cut down to one script. It expects Markdown files with a metadata header; for plain text, delete the line that drops the header.
# pip install tiktoken==0.14.0
import re, tiktoken
def body(path):
text = open(path, encoding="utf-8").read()
text = re.split(r"^---\s*$", text, maxsplit=2, flags=re.M)[2] # drop the header
return re.sub(r"\]\([^)]*\)", "]", text) # keep link text, drop URLs
pairs = [("en/page.md", "nl/pagina.md")] # add your own English/Dutch pairs
for name in ["o200k_base", "cl100k_base"]:
enc = tiktoken.get_encoding(name)
en = sum(len(enc.encode(body(e))) for e, _ in pairs)
nl = sum(len(enc.encode(body(n))) for _, n in pairs)
print(name, en, nl, round(nl / en, 3))
- Install the counter with
pip install tiktoken==0.14.0. (2 minutes) - Put your pairs in the list and run the script. It prints the English tokens, the Dutch tokens and the ratio for both encodings. (8 minutes)
- Write the ratio next to every AI task you run in Dutch, and multiply that task’s token budget by it. (5 minutes)
If your ratio sits near 1.2, your documents behave like ours. If it is much higher, you have just found where part of your AI budget goes.
Cite this:Dutch costs 20% more GPT tokens than English. We measured it on 78 article pairs.Len P. van der Hof. https://lenvanderhof.com/en/blog/dutch-costs-more-tokens/ ·