Retrieval-augmented generation (RAG) is the usual way to have a language model answer questions from your own documents: you first look up the relevant chunks and then hand them to the model as context. Many teams see it as the safe variant, because the model is not trained on your documents. That is true, but it says nothing about where those documents go while the system is in use. This article walks the chain step by step and shows what goes over the line at each stage, so you can judge whether inference in a Dutch data centre solves something for you or merely feels different.
The misconception: only the question goes out
The picture many people have: the documents stay with us, the API only gets the question. In reality a RAG setup has four components, each generating its own traffic:
- Embedding model — turns text into vectors. During indexing it sees every chunk of every document; with every question it sees the question itself.
- Vector database — stores the vectors and, in practice, the text chunks next to them, because you need to return those later.
- Reranker — optional but increasingly common: receives the question plus a stack of candidate chunks and sorts them by relevance.
- Chat model — receives the question and the best chunks as context, and writes the answer.
Put any one of these behind an external API and, at every step where that component is involved, part of your corpus crosses the line. Ask questions about your documentation for a year and you will eventually have sent a large part of it to the chat model.
What goes over the line at each step
It helps to look at the traffic per phase, because the volumes differ a lot.
- Indexing — the largest volume. Every document is cut into chunks and every chunk goes to the embedding model. This is your entire corpus, in plain text, over the line.
- Asking a question — small. The question goes to the embedding model, the vector to the vector database. A vector is a row of numbers; you cannot read the question back out of it.
- Searching and reranking — medium. The vector database returns candidate chunks; the reranker receives all of them plus the question. This is text again.
- Generating the answer — the question plus the selected chunks go to the chat model as a single prompt. Text again, with every question.
The point teams most often miss is the vector database. It sounds like a bin of abstract numbers, but the text chunk is almost always stored right next to the vector. A managed vector database outside your walls is therefore a copy of your corpus outside your walls, with the search traffic on top.
Re-indexing: bandwidth and egress
You do not index once. You switch embedding models, change the chunk size, or your documents get updated. Every full re-index sends your whole corpus to the embedding model again. If your documents live with a hyperscaler and the embedding model somewhere else, you also pay egress on that, and with most cloud providers outbound traffic is the expensive direction.
If storage, embedding model and vector database are in the same data centre, re-indexing is traffic inside the rack or over a cross-connect. There is no egress and bandwidth is rarely the bottleneck; the embedding model itself is.
Everything in one data centre: what it does and does not solve
The setup that keeps the traffic inside is simple to describe: your own servers holding the documents and the vector database in colocation, the embedding model, reranker and chat model in the same building, and between them a cross-connect or a private VRF on the EVPN fabric. Not a single chunk touches the public internet.
Be honest about what that means, though. It answers the question of where your documents are and who could see them in transit. It does not answer who can reach the servers, what gets logged, or whether a model trains on your data; you settle those contractually, with a data processing agreement and arrangements about logging. And even inside one data centre you still switch TLS on, because a private network is no reason to leave traffic unencrypted.
Honest about latency
Latency is often cited as the second argument, and it deserves nuance. A cross-connect within one data centre typically sits under 0.3 ms round trip; Ede–Amsterdam is around 1 to 2 ms and the Netherlands–United States around 80 to 120 ms. Those are typical values, not measured by us.
Set that against generating one chat answer: hundreds of milliseconds to seconds. For a user asking a single question the network gain disappears into the noise. Where it does count:
- Agents and multi-step RAG — a task that makes 10 to 100 calls in sequence pays the network cost on every round. A hundred times 100 ms is ten seconds; a hundred times 0.3 ms you will not notice.
- Indexing in batches — thousands of embedding calls in quick succession.
- Speech — transcription and answer in the same pipeline, where every round of waiting is audible.
So do not choose a single data centre because it feels faster, but because you know what kind of workload you run.
RAG at Xyphen IT
We provide AI inference in BIT-2C in Ede with an OpenAI-compatible API, shared per token or on dedicated cards. Embeddings, reranking and chat run there side by side; how you connect to them is on the technical page. If your documents and vector database sit in your own rack with us in BIT colocation, we link them with a cross-connect or a private VRF on our EVPN fabric. No training on your data and no storage of prompts are laid down in an addendum and the data processing agreement. If you do not have a RAG pipeline yet, we can build one as software development.
Want to know whether this fits your corpus and workload? Request a proposal or get in touch.