Knowledge baseAI

Choosing open models: licence, origin and Dutch

15 September 20267 min read

Anyone picking an open language model for the first time lands on a list of dozens of names and one reassuring thought: it is open, so it is allowed. That thought is wrong. In this field, "open" only says that you can download the weights. What you may do with them, who made them and how well they handle Dutch are three separate questions, and together they decide whether a model belongs in your environment. This article walks through them for the models we meet on our own inference platform.

Open-weight is not open source

A model consists of weights: a large file of numbers. When the maker publishes that file, the model is called open-weight. That says nothing about the training data, the training code or the terms. The terms are in the licence, and that is where the difference lies.

Apache 2.0 and MIT are genuine open-source licences: you may use the model commercially, modify it, fine-tune it and redistribute it without asking permission. A maker's own licence can deviate from that on any point. So do not read the model's name; read the licence next to it.

This is what that looks like right now for the best-known open models:

  • Mistral Small 3.2 (24B, France) — Apache 2.0. The larger Mistral models exist too, but need room that a mid-sized environment does not have.
  • gpt-oss-20b and 120b (OpenAI, US) — Apache 2.0.
  • EuroLLM-9B and 22B (EU consortium) — Apache 2.0, trained on all EU languages.
  • Gemma 3 27B (Google, US) — the Gemma terms of use, not an open-source licence. Usable, but you sign with the maker.
  • Qwen 3.x (Alibaba, China) — Apache 2.0. DeepSeek (China) — MIT.
  • Llama 4 (Meta, US) — the Llama Community License. Not an open-source licence, and for the multimodal Llama models (3.2 and later, and Llama 4) companies established in the EU are excluded from the licence altogether.

That last line is easily overlooked. A Dutch company deploying a multimodal Llama model does so without a valid licence, however open the model looks.

How good is it at Dutch

Most open models are trained predominantly on English. Dutch comes along on the side, and how well varies between models far more than the parameter counts suggest.

Two things to know. First, the benchmarks: on EuroEval, the European benchmark suite, the Gemma and Qwen models rank among the best open models for Dutch. EuroLLM is the only one in this list that treats Dutch as a first-class language; EuroLLM-9B scores well at rewriting into plain language, which for government and healthcare is often the main use. Mistral Small is reasonable in Dutch, not outstanding.

Second, the tokeniser. Dutch costs more tokens per word than English: roughly 1.3 to 1.6 tokens per word against around 1.3 for English, depending on the tokeniser. A Dutch prompt of 1,000 words can therefore be 1,600 tokens where the same text in English costs 1,300. That is 20 per cent more context used up and 20 per cent more to pay, on a model billed per token. It is an indication rather than a fixed factor, but count it in when comparing models.

Origin is a choice, not a detail

Technically, Qwen and DeepSeek are strong, and their licences are generous. Yet most of our customers do not pick them, and that has little to do with quality.

A model you run locally sends nothing to its maker. The weights are a file; there is no connection to China or the US inside them. That is different from using the same maker's app or API: Dutch civil servants are not allowed to use the DeepSeek app, for example, and that rule is about the service, not the weights. But reputation and dependency do matter. Anyone who has to explain to a municipality or a care provider which model writes the letters wants to answer that question without a footnote. And anyone who tunes their prompts to one model wants to know whether the maker will still release a successor under the same terms in two years.

That is why we put an origin label on every model: who made it and under which licence. Models from China we run on request only, never as the default choice. That is not a judgement on the engineering; it is so that the choice is yours and stays explainable.

Does it fit in memory

The last question is the most down-to-earth: does the model fit on the hardware you have. Memory decides that, not compute. A model that does not fit in memory does not run, however fast the card.

The rule of thumb for the weights is parameters times bytes per parameter:

  • BF16 — 2 bytes per parameter. A 24B-parameter model then needs about 48 GB.
  • 4-bit quantisation (AWQ or GPTQ) — about half a byte per parameter plus overhead. That same 24B model then fits in about 14 GB.
  • FP8 — 1 byte per parameter, only on newer hardware.

Quantisation costs a little quality, usually barely noticeable, and makes the difference between a model that fits on one card and one that needs two.

On top of that comes the KV cache: the working memory of every conversation running at that moment. It grows with the number of concurrent users and with the length of their context. So the weights decide whether a model fits; the KV cache decides how many users can be on it at once. Whoever only counts the weights ends up with a model that fits exactly and then serves one user at a time.

Open models at Xyphen IT

On our inference platform in Ede we run small and mid-sized open models, each with an origin label and a licence we can explain. Which models those are and how you address them through the API is on the technical page. If you want to run a model dedicated, or your own fine-tune on your own hardware next to our network, you can do that through GPU colocation in the same data centre as our BIT colocation.

Not sure which model suits your application? Request a proposal or get in touch, and we will work out the memory and the Dutch with you.

Frequently asked questions

Frequently asked questions

Is an open-weight model the same as an open-source model?

No. Open-weight means you can download the weights and run them yourself. Whether you may also use them commercially, modify them and redistribute them is set by the licence. Apache 2.0 and MIT are open-source licences; the Llama Community License and the Gemma terms of use are not, even though the weights are freely downloadable.

Which open model is best at Dutch?

On EuroEval, the Gemma and Qwen models rank among the best open models for Dutch. EuroLLM is trained on all EU languages and does well at rewriting into plain language. Mistral Small is reasonable in Dutch. Always test with your own texts; a benchmark says little about your domain.

How much memory does a 24-billion-parameter model need?

Multiply parameters by bytes per parameter. In BF16 (2 bytes) that is about 48 GB; with 4-bit quantisation about 14 GB including overhead. On top of that comes the KV cache, which grows with the number of concurrent users and the length of their context.

Answer not found?

Ask an engineer directly — we usually respond within one business day.