AI inference is the moment a trained model does something for you: you send a prompt, the model computes, text comes back. It is the only part of AI you touch every day, and yet in conversations about "where is my data" it is routinely lumped together with training. That difference is exactly what this article is about, and it connects to the question we always ask about colocation too: which building is it in, and who holds the key?
Training and inference are two different things
Training is the one-off job of building a model: weeks of computation over a large body of text, resulting in a set of weights. That is a file. A model with 24 billion parameters is roughly 48 GB at 16-bit precision, and roughly 14 GB quantised to 4 bits. Open models are published as exactly such a file, and whoever downloads it has the same model as everyone else.
Inference is using that file. It is loaded into GPU memory, and every prompt that arrives is pushed through those weights. Nothing about the model changes; your prompt teaches it nothing. The only thing of yours that enters the machine is the prompt itself, and the only thing that leaves is the answer.
The misconception is that "using AI" means your data disappears into a model. That only happens if someone keeps your prompts and trains on them later. Whether that happens is not a property of the technology but of the agreement, so that is what you read.
Every call is a data transfer
What the technology does enforce: the model can only compute on text it has in front of it in readable form. A prompt is therefore never a bit of metadata but the full content. For a summary of a customer call, the whole call goes along. In a RAG setup, the retrieved document fragments go along. In a chat with history, the entire history goes along again on every turn.
That makes inference different from most SaaS. With an accounting package you send structured fields; with a language model you send prose, and prose contains names, diagnoses and contract terms. The Dutch Data Protection Authority warned, not without reason, about data breaches caused by employees pasting patient or customer data into public chatbots. Whoever hosts the model is legally a processor, and with this kind of content a DPIA is often required.
TLS protects the road, not the destination
Here is the second misconception: "the API call runs over HTTPS, so it is encrypted." That is true, and it is relevant. But TLS ends at the receiver. There your prompt is decrypted, because a model cannot compute on encrypted text. From that moment the content sits in a server's working memory, and in the logs if those are on.
What happens with it next is determined not by the encryption but by three other things:
- The contract — is the prompt kept, for how long, and is it trained on? That belongs in a data processing agreement, not in a privacy statement that can change unilaterally.
- The ownership chain — the US CLOUD Act applies to every US company and its subsidiaries that has data in its "possession, custody or control", regardless of where that data is stored. An EU region is therefore not EU jurisdiction. In June 2025, Microsoft France stated under oath before the French Senate that it cannot guarantee that French data will not be handed over to US authorities.
- The physical location — which law applies, who can physically get to it, and whether you can go there yourself.
To be fair: Dutch and European authorities can request data too, under the Dutch Intelligence and Security Services Act, criminal law or the European e-Evidence rules. The difference is that you then know under which law that happens, and which court you can go to.
What "where it runs" means in practice
To a network person, the model's location is just a hop in the path. A model in a data centre on the other side of the ocean is 80 to 120 ms away; a model in the same building, over a cross-connect, less than a millisecond. Those are typical values, not measurements of ours.
For a single chat answer that makes no difference. Generation itself takes hundreds of milliseconds to seconds, and 100 ms disappears against that. The difference appears in work that consists of many calls: an agent that consults the model ten to a hundred times per task, a RAG pipeline that first creates embeddings, then searches, then reranks and only then generates, or a nightly batch over a document archive. There every network round trip counts, and there "is the model next to my data" suddenly becomes a turnaround-time question.
And then there is the path itself. A model over the public internet has a public endpoint and an API key that works from the internet; a model behind a cross-connect or private VRF does not. That is not a compliance argument, but a smaller attack surface.
AI inference at Xyphen IT
We run open models in our rack in BIT-2C in Ede, reachable through an OpenAI-compatible API. How that API works and which network paths exist, from TLS over the internet to a cross-connect, is on the technical page; the decision itself is on AI inference. If you want your application or vector database next to the model, you can put it in colocation at BIT with a private VRF on our EVPN fabric, so the prompt never leaves the building. No training on your data and no storage of prompts is laid down in an addendum to your contract, with a data processing agreement. Want to know whether this fits what you are building? Request a proposal or get in touch.