AI inference that stays in Ede.
Open AI models on GPUs in the BIT-2C data centre, reachable through an OpenAI-compatible API.
You run open models on cards in our rack in Ede and call them the way you call a cloud API today. The difference is in what does not happen: we do not store your prompts and do not train on them, laid down in your contract, and processing takes place only in Ede. Put your own servers next to them and your data need not even leave the building.
Every prompt is a data transfer.
An AI call feels like a function call, but it is a data transfer. Every prompt contains whatever you are showing the model at that moment: a customer file, a contract, a piece of source code, the transcript of a conversation. With a cloud API that leaves the building, to a party whose terms you have accepted but whose processing you cannot check.
The misconception is that TLS solves this. TLS protects the road, not the destination. What happens to your prompt on the other side, how long it is kept and under which law, is in a policy document you did not negotiate. With us the model is in Ede and the arrangements are in your own contract.
AI inference with us fits if you:
- Want to run documents, files or source code through a model and do not want to send them to a cloud API outside the EU
- Are building a RAG or document pipeline and would rather put the vector database next to the model than behind it on the internet
- Must be able to explain under the GDPR, NIS2, BIO2 or DORA where processing takes place and who has access
- Already have an OpenAI-compatible integration and only want to change the base URL
- Want predictable costs for a fixed workload, without counting per token
One endpoint, one rack, one contract.
You get a working endpoint and the arrangements around it. No console with a hundred buttons, but an engineer you can talk to.
What we do not deliver: frontier models. What we do deliver is the right open model for your task, on hardware in Ede instead of in a worldwide queue.
The model in the same building as your data.
This is the combination we deliver as one package: colocation in the same data centre as the inference. Your servers, your vector database and your application sit next to the model, connected with a cross-connect or a private VRF on our EVPN network.
- Every prompt and every document goes over the public internet to an endpoint you do not manage
- An API key that works from any address on the internet
- Egress costs and internet bandwidth for every batch of documents
- Processing in a region of the provider's choosing, under its terms
- Model, vector database and application in the same building; the traffic does not go over the internet
- No public endpoint and no API key on the internet: the endpoint is reachable only from your own VRF
- No egress costs, bandwidth at LAN speed for RAG and document pipelines
- Reachable from your other locations over our own network, or over the internet with TLS or WireGuard
Hybrid works too: base load on your own GPU servers in colocation, peaks on the shared capacity, the same API.
Let the model look where the problem is.
Through MCP, an open standard for tools an AI model may use, the model can investigate in your own environment with your question as its brief. Why does that VPN tunnel keep flapping, which rule blocks this traffic, which device is filling the line. Connecting a FortiGate or a UniFi Dream Machine Pro is routine work for us.
- 01Read-onlyBy default the model only looks
Configuration, logs, sessions and statistics. A change only happens after a person approves it, and only if that was agreed beforehand.
- 02PrivateOver your own connection
Everything between the model and your environment can run over a cross-connect or private VRF. Your firewall does not need to expose a management port on the internet.
- 03Scoped accountYour key, your limits
The model works with its own read-only account on your equipment, which you can revoke yourself. What was requested and when, you record in your own logging.
- 04CustomYour own systems too
FortiGate and UniFi are examples. If your equipment or software has an API, we build the connector in consultation.
An AI model can be wrong. That is why it investigates and advises, and a person decides.
Shared by usage, dedicated per month.
One endpoint with a fixed model selection; you pay by usage. For those starting out, testing or with a variable workload. The same arrangements on data and processing as with dedicated.
Fixed cards for you alone, at a fixed monthly amount. You choose the model, including your own open weights if they fit on the hardware. Nobody else shares your capacity, your latency or your queue.
The right open model for your task.
Small and medium-sized open models, up to roughly 30 billion parameters per model; larger by arrangement. For most tasks in an organisation that is plenty, and for every model you know who made it and under which licence you use it. Below is a sample selection; we agree the current line-up with you.
General assistant, tool calling
Apache 2.0Strong in Dutch
Gemma Terms of UseFast, reasoning
Apache 2.0All EU languages, plain language
Apache 2.0Embeddings for RAG, multilingual
MITSpeech to text
MITChinese language models, such as Qwen, we deploy only on request, with an origin label, never by default. Your own open weights run on dedicated cards if they fit. Which models are available at the moment, we agree during the intake.
Measured consumption, not an estimate.
On request you receive a monthly report with the measured power consumption of your inference, attributed to what you used, and the CO2 that goes with it. Usable for your VSME or CSRD reporting under scope 3.
- 01Power in kWh, measured
The consumption of the GPUs and the server itself, from the meter, not from a datasheet.
- 02Attributed to your usage
For shared by your tokens, for dedicated by your cards. You see what your share was, not what the rack did.
- 03The data centre's share
Cooling and losses via the PUE of our hall, measured live by BIT, so the figure matches that month and not an annual average.
- 04CO2 by two methods
Location-based and market-based, with source and emission factor included. BIT takes Dutch power from sun and wind, demonstrated with cancelled Guarantees of Origin.
The report is optional and on request. We do not call the inference green or climate neutral; you get the measurement and the sources, and judge for yourself.
From intake to endpoint.
- IntakeWhich task, which data, how much traffic. That leads to a proposal: shared, dedicated or a combination, with the model we recommend and why.
- Contract and addendumData processing agreement and the addendum on training and storage. Only once that is signed does production data come anywhere near a card.
- Endpoint and accessYou get your endpoint: over the internet with TLS, through WireGuard, or through a cross-connect or private VRF if your servers are next to it. We walk through the first calls with you.
- In useYou see your usage, we keep the models and the hardware up to date. Updates go in a maintenance window we announce in advance.
You can try a model first on the shared capacity, with test data, before connecting anything from production.


Network, data centre and model from a single source.
- Our own network, AS211588, with core locations at NorthC and Nikhef and rack space in all BIT data centres in Ede
- Colocation, cross-connects and private VRFs are our daily work; the model is the new neighbour in the rack
- Engineers you talk to directly, including about quantisation, context and latency
- Arrangements on your data are in the contract, not on a policy page
We do not compete on price. We are the party that understands the model, the network and the rack around it.
Frequently asked questions about AI inference
Are these models good enough compared with the big cloud models?
For many tasks in an organisation, yes: summarising, classifying, extraction, question answering over your own documents, transcription. A model of 24 or 27 billion parameters does that well, certainly with RAG on your own sources. For open-ended reasoning about anything and everything a frontier model remains stronger; we do not deliver that. In the intake we tell you honestly whether your task fits.
What about the GDPR and the data processing agreement?
As a rule we are the processor and you are the controller. The data processing agreement lays down that processing takes place only in Ede, and the addendum that we do not train on your data and do not store your prompts. That helps with your accountability under the GDPR, NIS2, BIO2, DORA and the AI Act. It does not make you compliant by itself; that remains your own work.
How do I switch from the OpenAI API?
You change the base URL and the key in your client and pick a model name from our selection. Chat completions, embeddings and transcription work through the same paths. Prompts tuned to a large model usually deserve a once-over; we look at that with you in the first week.
Do you also run Chinese models?
Only on request, with an origin label. Models such as Qwen are technically good, but we never put them in the shared selection by default. If you want one on dedicated cards, that is possible, and it is labelled with where it comes from.
How much energy does this use?
We measure it. On request you receive a monthly report with kWh, attributed to your usage, the data centre's share and the CO2 by the location-based and market-based methods, with source and emission factor. No green labels: you get the measurement.
What if my model does not fit?
Models up to roughly 30 billion parameters fit on our cards, often also in a quantised variant with little loss of quality. Larger is by arrangement: sometimes it fits with quantisation, sometimes the answer is no. If it does not fit, we say so up front, not after the first invoice.
Want to know which model fits your task?
Tell us what you want to run through the model and how much. After an intake you receive a proposal with the model, the form and the arrangements on your data.