ServicesAI

AI inference that stays in Ede.

Open AI models on GPUs in the BIT-2C data centre, reachable through an OpenAI-compatible API.

You run open models on cards in our rack in Ede and call them the way you call a cloud API today. The difference is in what does not happen: we do not store your prompts and do not train on them, laid down in your contract, and processing takes place only in Ede. Put your own servers next to them and your data need not even leave the building.

OpenAI-compatible APIProcessing in BIT-2C, EdeShared or dedicatedYour colocation alongside

Every prompt is a data transfer.

An AI call feels like a function call, but it is a data transfer. Every prompt contains whatever you are showing the model at that moment: a customer file, a contract, a piece of source code, the transcript of a conversation. With a cloud API that leaves the building, to a party whose terms you have accepted but whose processing you cannot check.

The misconception is that TLS solves this. TLS protects the road, not the destination. What happens to your prompt on the other side, how long it is kept and under which law, is in a policy document you did not negotiate. With us the model is in Ede and the arrangements are in your own contract.

Who it is for

AI inference with us fits if you:

  • Want to run documents, files or source code through a model and do not want to send them to a cloud API outside the EU
  • Are building a RAG or document pipeline and would rather put the vector database next to the model than behind it on the internet
  • Must be able to explain under the GDPR, NIS2, BIO2 or DORA where processing takes place and who has access
  • Already have an OpenAI-compatible integration and only want to change the base URL
  • Want predictable costs for a fixed workload, without counting per token
What you get

One endpoint, one rack, one contract.

You get a working endpoint and the arrangements around it. No console with a hundred buttons, but an engineer you can talk to.

OpenAI-compatible API
Chat, embeddings and transcription through the same paths your client library already knows. Change the base URL and the key; the rest of your code stays as it is.
Processing in Ede, and nowhere else
The model runs in our rack in BIT-2C. No failover to another region and no American company in the chain: Xyphen IT, BIT and the owner of the hardware are Dutch.
No training, no storage of prompts
Laid down in an addendum to your contract and confirmed in writing in the data processing agreement. Not a promise on a website, but an arrangement in your contract.
No lock-in
Open weights, a standard API and no exit fees. If it fits better elsewhere, you take your prompts and your pipeline with you.

What we do not deliver: frontier models. What we do deliver is the right open model for your task, on hardware in Ede instead of in a worldwide queue.

Next to the model

The model in the same building as your data.

This is the combination we deliver as one package: colocation in the same data centre as the inference. Your servers, your vector database and your application sit next to the model, connected with a cross-connect or a private VRF on our EVPN network.

Through a cloud API
  • Every prompt and every document goes over the public internet to an endpoint you do not manage
  • An API key that works from any address on the internet
  • Egress costs and internet bandwidth for every batch of documents
  • Processing in a region of the provider's choosing, under its terms
Next to the model at BIT
  • Model, vector database and application in the same building; the traffic does not go over the internet
  • No public endpoint and no API key on the internet: the endpoint is reachable only from your own VRF
  • No egress costs, bandwidth at LAN speed for RAG and document pipelines
  • Reachable from your other locations over our own network, or over the internet with TLS or WireGuard

Hybrid works too: base load on your own GPU servers in colocation, peaks on the shared capacity, the same API.

In your own environment

Let the model look where the problem is.

Through MCP, an open standard for tools an AI model may use, the model can investigate in your own environment with your question as its brief. Why does that VPN tunnel keep flapping, which rule blocks this traffic, which device is filling the line. Connecting a FortiGate or a UniFi Dream Machine Pro is routine work for us.

  1. 01Read-only
    By default the model only looks

    Configuration, logs, sessions and statistics. A change only happens after a person approves it, and only if that was agreed beforehand.

  2. 02Private
    Over your own connection

    Everything between the model and your environment can run over a cross-connect or private VRF. Your firewall does not need to expose a management port on the internet.

  3. 03Scoped account
    Your key, your limits

    The model works with its own read-only account on your equipment, which you can revoke yourself. What was requested and when, you record in your own logging.

  4. 04Custom
    Your own systems too

    FortiGate and UniFi are examples. If your equipment or software has an API, we build the connector in consultation.

An AI model can be wrong. That is why it investigates and advises, and a person decides.

Two forms

Shared by usage, dedicated per month.

Per token
Shared

One endpoint with a fixed model selection; you pay by usage. For those starting out, testing or with a variable workload. The same arrangements on data and processing as with dedicated.

Per month
Dedicated

Fixed cards for you alone, at a fixed monthly amount. You choose the model, including your own open weights if they fit on the hardware. Nobody else shares your capacity, your latency or your queue.

Set out your use case in the quote tool →
Models

The right open model for your task.

Small and medium-sized open models, up to roughly 30 billion parameters per model; larger by arrangement. For most tasks in an organisation that is plenty, and for every model you know who made it and under which licence you use it. Below is a sample selection; we agree the current line-up with you.

Mistral Small 3.2 (24B)
Mistral AI · France

General assistant, tool calling

Apache 2.0
Gemma 3 (27B)
Google · United States

Strong in Dutch

Gemma Terms of Use
gpt-oss-20b
OpenAI · United States

Fast, reasoning

Apache 2.0
EuroLLM-9B
European consortium · EU

All EU languages, plain language

Apache 2.0
multilingual-e5-large
Microsoft · United States

Embeddings for RAG, multilingual

MIT
Whisper large-v3
OpenAI · United States

Speech to text

MIT

Chinese language models, such as Qwen, we deploy only on request, with an origin label, never by default. Your own open weights run on dedicated cards if they fit. Which models are available at the moment, we agree during the intake.

Energy and CO2

Measured consumption, not an estimate.

On request you receive a monthly report with the measured power consumption of your inference, attributed to what you used, and the CO2 that goes with it. Usable for your VSME or CSRD reporting under scope 3.

  1. 01Power in kWh, measured

    The consumption of the GPUs and the server itself, from the meter, not from a datasheet.

  2. 02Attributed to your usage

    For shared by your tokens, for dedicated by your cards. You see what your share was, not what the rack did.

  3. 03The data centre's share

    Cooling and losses via the PUE of our hall, measured live by BIT, so the figure matches that month and not an annual average.

  4. 04CO2 by two methods

    Location-based and market-based, with source and emission factor included. BIT takes Dutch power from sun and wind, demonstrated with cancelled Guarantees of Origin.

The report is optional and on request. We do not call the inference green or climate neutral; you get the measurement and the sources, and judge for yourself.

How we work

From intake to endpoint.

  1. Intake
    Which task, which data, how much traffic. That leads to a proposal: shared, dedicated or a combination, with the model we recommend and why.
  2. Contract and addendum
    Data processing agreement and the addendum on training and storage. Only once that is signed does production data come anywhere near a card.
  3. Endpoint and access
    You get your endpoint: over the internet with TLS, through WireGuard, or through a cross-connect or private VRF if your servers are next to it. We walk through the first calls with you.
  4. In use
    You see your usage, we keep the models and the hardware up to date. Updates go in a maintenance window we announce in advance.

You can try a model first on the shared capacity, with test data, before connecting anything from production.

In the BIT data centre
The cold aisle in the BIT data centre
The cold aisle in the BIT data centre
The hot aisle in the BIT data centre
The hot aisle in the BIT data centre
Why Xyphen IT

Network, data centre and model from a single source.

  • Our own network, AS211588, with core locations at NorthC and Nikhef and rack space in all BIT data centres in Ede
  • Colocation, cross-connects and private VRFs are our daily work; the model is the new neighbour in the rack
  • Engineers you talk to directly, including about quantisation, context and latency
  • Arrangements on your data are in the contract, not on a policy page

We do not compete on price. We are the party that understands the model, the network and the rack around it.

FAQ

Frequently asked questions about AI inference

Are these models good enough compared with the big cloud models?

For many tasks in an organisation, yes: summarising, classifying, extraction, question answering over your own documents, transcription. A model of 24 or 27 billion parameters does that well, certainly with RAG on your own sources. For open-ended reasoning about anything and everything a frontier model remains stronger; we do not deliver that. In the intake we tell you honestly whether your task fits.

What about the GDPR and the data processing agreement?

As a rule we are the processor and you are the controller. The data processing agreement lays down that processing takes place only in Ede, and the addendum that we do not train on your data and do not store your prompts. That helps with your accountability under the GDPR, NIS2, BIO2, DORA and the AI Act. It does not make you compliant by itself; that remains your own work.

How do I switch from the OpenAI API?

You change the base URL and the key in your client and pick a model name from our selection. Chat completions, embeddings and transcription work through the same paths. Prompts tuned to a large model usually deserve a once-over; we look at that with you in the first week.

Do you also run Chinese models?

Only on request, with an origin label. Models such as Qwen are technically good, but we never put them in the shared selection by default. If you want one on dedicated cards, that is possible, and it is labelled with where it comes from.

How much energy does this use?

We measure it. On request you receive a monthly report with kWh, attributed to your usage, the data centre's share and the CO2 by the location-based and market-based methods, with source and emission factor. No green labels: you get the measurement.

What if my model does not fit?

Models up to roughly 30 billion parameters fit on our cards, often also in a quantised variant with little loss of quality. Larger is by arrangement: sometimes it fits with quantisation, sometimes the answer is no. If it does not fit, we say so up front, not after the first invoice.

Want to know which model fits your task?

Tell us what you want to run through the model and how much. After an intake you receive a proposal with the model, the form and the arrangements on your data.