Knowledge baseAI

What is AI inference, and why it matters where it runs

11 September 20266 min read

AI inference is the moment a trained model does something for you: you send a prompt, the model computes, text comes back. It is the only part of AI you touch every day, and yet in conversations about "where is my data" it is routinely lumped together with training. That difference is exactly what this article is about, and it connects to the question we always ask about colocation too: which building is it in, and who holds the key?

Training and inference are two different things

Training is the one-off job of building a model: weeks of computation over a large body of text, resulting in a set of weights. That is a file. A model with 24 billion parameters is roughly 48 GB at 16-bit precision, and roughly 14 GB quantised to 4 bits. Open models are published as exactly such a file, and whoever downloads it has the same model as everyone else.

Inference is using that file. It is loaded into GPU memory, and every prompt that arrives is pushed through those weights. Nothing about the model changes; your prompt teaches it nothing. The only thing of yours that enters the machine is the prompt itself, and the only thing that leaves is the answer.

The misconception is that "using AI" means your data disappears into a model. That only happens if someone keeps your prompts and trains on them later. Whether that happens is not a property of the technology but of the agreement, so that is what you read.

Every call is a data transfer

What the technology does enforce: the model can only compute on text it has in front of it in readable form. A prompt is therefore never a bit of metadata but the full content. For a summary of a customer call, the whole call goes along. In a RAG setup, the retrieved document fragments go along. In a chat with history, the entire history goes along again on every turn.

That makes inference different from most SaaS. With an accounting package you send structured fields; with a language model you send prose, and prose contains names, diagnoses and contract terms. The Dutch Data Protection Authority warned, not without reason, about data breaches caused by employees pasting patient or customer data into public chatbots. Whoever hosts the model is legally a processor, and with this kind of content a DPIA is often required.

TLS protects the road, not the destination

Here is the second misconception: "the API call runs over HTTPS, so it is encrypted." That is true, and it is relevant. But TLS ends at the receiver. There your prompt is decrypted, because a model cannot compute on encrypted text. From that moment the content sits in a server's working memory, and in the logs if those are on.

What happens with it next is determined not by the encryption but by three other things:

  • The contract — is the prompt kept, for how long, and is it trained on? That belongs in a data processing agreement, not in a privacy statement that can change unilaterally.
  • The ownership chain — the US CLOUD Act applies to every US company and its subsidiaries that has data in its "possession, custody or control", regardless of where that data is stored. An EU region is therefore not EU jurisdiction. In June 2025, Microsoft France stated under oath before the French Senate that it cannot guarantee that French data will not be handed over to US authorities.
  • The physical location — which law applies, who can physically get to it, and whether you can go there yourself.

To be fair: Dutch and European authorities can request data too, under the Dutch Intelligence and Security Services Act, criminal law or the European e-Evidence rules. The difference is that you then know under which law that happens, and which court you can go to.

What "where it runs" means in practice

To a network person, the model's location is just a hop in the path. A model in a data centre on the other side of the ocean is 80 to 120 ms away; a model in the same building, over a cross-connect, less than a millisecond. Those are typical values, not measurements of ours.

For a single chat answer that makes no difference. Generation itself takes hundreds of milliseconds to seconds, and 100 ms disappears against that. The difference appears in work that consists of many calls: an agent that consults the model ten to a hundred times per task, a RAG pipeline that first creates embeddings, then searches, then reranks and only then generates, or a nightly batch over a document archive. There every network round trip counts, and there "is the model next to my data" suddenly becomes a turnaround-time question.

And then there is the path itself. A model over the public internet has a public endpoint and an API key that works from the internet; a model behind a cross-connect or private VRF does not. That is not a compliance argument, but a smaller attack surface.

AI inference at Xyphen IT

We run open models in our rack in BIT-2C in Ede, reachable through an OpenAI-compatible API. How that API works and which network paths exist, from TLS over the internet to a cross-connect, is on the technical page; the decision itself is on AI inference. If you want your application or vector database next to the model, you can put it in colocation at BIT with a private VRF on our EVPN fabric, so the prompt never leaves the building. No training on your data and no storage of prompts is laid down in an addendum to your contract, with a data processing agreement. Want to know whether this fits what you are building? Request a proposal or get in touch.

Frequently asked questions

Frequently asked questions

Is inference the same as training a model?

No. Training is the one-off job of building the model from a large body of data; inference is using it, every time you send a prompt. An open model you run yourself has already been trained. Your prompts are not used for that unless someone explicitly does so.

If the connection is encrypted with TLS, is my prompt safe?

In transit, yes. TLS protects the transport, but at the receiving end your prompt is decrypted, because otherwise the model cannot do anything with it. What happens after that, how long it is kept and who can access it is not decided by the encryption but by the contract and by where the hardware is.

Do I notice the distance to the model?

For a single chat answer, hardly. Generation itself takes hundreds of milliseconds to seconds, and a network delay of tens of milliseconds disappears against that. It does add up for agents and RAG pipelines that make dozens of calls per task, and for embedding batches.

Answer not found?

Ask an engineer directly — we usually respond within one business day.