Tokens are the basic units of data for LLMs, similar to how words form the basis of sentences. A token can take different forms, though. It might be part of a word, an entire word, or even punctuation. LLMs work as next-token predictors, calculating the probability of the next token based on the sequence of preceding tokens.
Let’s take token itself as an example. “Tokens” is 1 token, but “tokenize” can be 2.
- This sentence is 7 tokens.
- We need to tokenize the text for better analysis.
Whether "tokenize" is one or two tokens depends on how often the word appears in the training set of the Large Language Model. In models trained on large vocabularies, where there is less need to split familiar words into smaller units, "tokenize" would typically be a single token. Where the vocabulary is smaller or the word is less common, it may be split into two tokens.
Token usage differences across languages
The performance of Large Language Models varies across languages. We won't go into the specifics here, but the number of tokens needed for a typical sentence differs a lot between languages, and that is worth knowing.
English, with thousands of words represented as single tokens, is usually the most economical in token usage. This is partly due to how the LLM encodes text, which tends to favour English because of its dominant presence in the training data. As a result, English is a cheap language for LLM operations.
For instance, consider the word "probability", a relatively long word in English, yet it counts as only one token. Translated into other languages, the token count is different;
- NL - Kans - 2 tokens
- DE - Wahrscheinlichkeit - 4 tokens
- FR - Probabilité - 2 tokens
- ES - Probabilidad - 2 tokens
An exact comparison is hard, but when it comes to token usage, major European languages like German, French, Spanish and Dutch (not a large language, we know) are roughly 1.5 to 1.7 times more token-intensive than English. Smaller European languages can be even more demanding, with Hungarian at around 2.5 times and Greek at about 4 times the token usage of English.
For a major language like Hindi, token usage is around 5 times that of English. That is not only due to the complexity of the language but also to its script rules and the use of conjunct characters.

Why is this important?
Since we started from the premise that token usage equals cost, languages that need more tokens to say the same thing can be considerably more expensive.
The context window is closely related. The context window is the amount of text (measured in tokens) that a model can process at once when generating a response or understanding a text. This means that working in a language like Hindi not only increases operational costs but also limits the amount of context you can include. We'll look at context windows in more detail in an upcoming blog post.
While major players like OpenAI and Anthropic charge per token, Google's PaLM2 model has introduced charging per character, which could reduce costs under certain conditions. Another alternative is deploying an open-source model, typically billed per hour of use. However, this option often involves higher setup and fixed monthly fees, making it an optimisation rather than a launch strategy. Moreover, the quality of open-source models may not meet the requirements for certain use cases.
To find the most suitable option for your needs, have a look at our design sprints. We will review your content and help you identify the best approach for your specific situation.

