Knowledge baseAI

Tokens explained: what an LLM call really costs

22 September 20266 min read

With a language-model API you pay per token, and most discussions about what AI costs run aground on that word. A token is not a word and not a character, but a piece of text as the model's tokenizer cuts it: a short word is often one token, a longer word two or three. If you want to know what a call costs, you first need to know how many tokens go in and come out. That builds on what inference is: every call sends the full content along, and that content is what you are billed for.

A worked example with a Dutch prompt

Take an everyday instruction, in Dutch: "Vat dit klantgesprek samen in drie punten en stel een vervolgafspraak voor." ("Summarise this customer call in three points and propose a follow-up appointment.") That is twelve words. As an indication, Dutch costs 1.3 to 1.6 tokens per word, so this sentence alone is around 16 to 19 tokens. That is not where the money is.

The customer call itself goes along. A transcript of 800 words is 1,000 to 1,300 tokens. On top of that comes the system prompt in which you tell the model who it is and how to answer, say 100 words, another 130 to 160 tokens. The input of this single call therefore lands around 1,200 to 1,500 tokens. The answer, three points and a proposal, is perhaps 150 words: 200 to 240 tokens of output.

The ratio is thus roughly six to one. That is typical: in most applications the input is many times larger than the output, because you send context along and get a short answer back.

Input and output have two prices

Providers charge a different price for input tokens than for output tokens, and that is not marketing but engineering. The input goes through the model in a single pass, the prefill. The output is produced token by token, the decode, and every next token has to go through all the weights again. Per token, output therefore costs far more compute time than input.

To get a feel for the order of magnitude (market indication, not a price of ours): a European provider charges around EUR 0.15 per million input tokens and EUR 0.35 per million output tokens for a mid-sized open model. For the example above, 1,400 tokens in and 220 out, that is around EUR 0.0003 per call. At 10,000 such calls per day you arrive at about EUR 3 per day, some EUR 80 to 90 per month.

That is little, and that is also the point: per token is almost always the cheapest start. The costs only climb because of what you do not yet see in this sum.

Where the sum goes wrong

  • Chat history — a model has no memory between calls. On every turn the full history goes along again. A conversation of twenty turns sends its first question twenty times, and the twentieth call carries the whole conversation as input. Long chats are therefore quadratically more expensive than the individual messages suggest.
  • RAG and agents — a pipeline that first searches, then sends five document fragments along and then generates easily has 3,000 to 6,000 input tokens per user question. An agent that consults the model ten to a hundred times per task multiplies that again.
  • Dutch — tokenizers are trained predominantly on English. A compound such as "vervolgafspraak" is cut into pieces where the English "follow-up" costs two. The same text costs, as an indication, ten to twenty per cent more tokens in Dutch than in English, and you pay that difference on every call.

The practical lesson: measure your tokens before you promise anything. Every OpenAI-compatible API returns, with each answer, how many input and output tokens the call cost. A week of logging says more than any estimate, just as with power a week of measuring says more than the nameplate.

Shared per token or dedicated per month

Per token you pay only for what you use, and you share the GPU with others. Dedicated means a card, or a set of cards, is yours: you pay a fixed amount per month whether it is computing or idle. As a market indication, a high-end data-centre card in Europe costs somewhere between EUR 2.70 and EUR 4.50 per hour, so roughly EUR 2,000 to 3,300 per month.

Put that next to the example of EUR 80 to 90 per month and the break-even is a long way off: you need to process twenty to forty times more before dedicated wins purely on cost. That is a continuously loaded card, not an office-hours pattern. If nothing happens at night and at the weekend, dedicated means paying for air.

Still, people opt for dedicated sooner than that, usually for reasons other than the token price:

  • Predictability — a fixed amount instead of an invoice that grows with a campaign or an enthusiastic department.
  • Isolation — no neighbours on the same card, so no fluctuating response times when someone else runs a batch.
  • Model choice — a model the shared service does not offer, or your own fine-tune.
  • Batch work — if you have the card anyway, you can process an archive overnight. Batching also makes inference considerably more frugal per token: in one published measurement the energy per token differed by roughly a factor of ten between small and large batches.

The honest answer to "when does it flip" is therefore: start per token, log your usage, and after a month work out whether a card of your own would be busy more than half the time. If not, keep sharing.

Tokens at Xyphen IT

Our AI inference is available both ways from day one: shared per token through an OpenAI-compatible API, or dedicated with fixed cards per month. How the API reports tokens and which models run is on the technical page. If you want your own GPU servers next to the model, the power and cooling side is covered under GPU colocation, and if you want a vector database or application in the same building, under colocation at BIT. Torn between shared and dedicated? Request a proposal and we will work it out with your actual tokens.

Frequently asked questions

Frequently asked questions

Why are output tokens more expensive than input tokens?

Because they are processed differently. The input goes through the model in one pass (prefill); the output is produced token by token (decode), and every step has to go through the weights again. Per token, output therefore costs far more compute time, which is why providers charge two prices.

In a chat, do I only pay for my new message?

No. A model has no memory between calls; on every turn the full history goes along again as input. A conversation of twenty turns therefore sends its first question twenty times. That is why the cost of long chats rises faster than the individual messages suggest.

Does Dutch really cost more tokens than English?

Yes, as an indication around ten to twenty per cent more per word, with differences between tokenizers. Tokenizers are trained mostly on English; Dutch compounds and inflections are therefore cut into pieces more often. The same text in Dutch costs more tokens than its English translation.

Answer not found?

Ask an engineer directly — we usually respond within one business day.