ServicesAITechnical

Inference, under the hood.

How the API, the cards and the network fit together.

This page is for whoever builds the integration. Which endpoints there are, how we quantise models, what dedicated isolates exactly, and the four routes your traffic can take to the model.

Chat, embeddings, transcriptionStreaming and tool callingNo logging of promptsTLS, WireGuard, cross-connect, VRFProcessing in BIT-2C, Ede

Compatible is not the same as identical.

Our API follows the OpenAI schema for chat completions, embeddings and audio transcription, with streaming and tool calling. Your existing client works after changing the base URL. What is not the same: the models behind it are open models up to roughly 30 billion parameters, the context is what the model can handle, and parameters that exist only for a specific cloud model are ignored. We would rather say that up front than have you discover it in production.

What you get
Endpoint and keys
One endpoint per environment and keys per application, which we replace or revoke at your request. With dedicated, an endpoint that exists only in your VRF, if you want.
Model selection with versions
Every model name points to a fixed version and quantisation. A new version gets a new name; you choose when to switch.
Usage and limits
Tokens per model and per key, and limits we set with you, so that a loop in a script does not become a surprise.
Arrangements on paper
No training, no storage of prompts: in the addendum to your contract and confirmed in writing in the data processing agreement.

The code to get started is below; the rest you discuss with an engineer, not a ticket form.

Hands-on

Change two lines, the rest stays as it is.

A chat completion against our endpoint, with streaming on. The only things that differ from a cloud API: the base URL and the model name.

Request
# Bestaande OpenAI-client, alleen de base_url wijzigt
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.example.nl/v1",
    api_key="sk-…",
)

reply = client.chat.completions.create(
    model="mistral-small-3.2",
    messages=[{"role": "user",
               "content": "Vat deze offerte samen in drie zinnen."}],
    stream=True,
)
Response
$ curl -s https://inference.example.nl/v1/models \
    -H "Authorization: Bearer $API_KEY" | jq -r '.data[].id'
mistral-small-3.2
gemma-3-27b
gpt-oss-20b
eurollm-9b
multilingual-e5-large
whisper-large-v3

$ curl -s https://inference.example.nl/v1/embeddings \
    -H "Authorization: Bearer $API_KEY" \
    -d '{"model": "multilingual-e5-large", "input": "Wat is colocatie?"}' \
    | jq '.data[0].embedding | length'
1024

Embeddings and transcription go through the same client; only the path and the model differ.

Under the hood

What sits between your request and the answer.

Seven things you want to know before you build on it.

API
Chat, embeddings, transcription, streaming

Chat completions with streaming, tool calling and JSON output, embeddings for RAG and audio transcription. The same paths and fields as the OpenAI schema; parameters that exist only for a specific cloud model are ignored.

Quantisation
Fewer bits, measured quality

Where it makes sense, models run in a quantised variant so that context and throughput fit on the card. For each model we state which variant it is. If you want full precision on dedicated, we work out what that costs in context and speed.

Isolation
Dedicated is really dedicated

With dedicated, your model runs on cards assigned to nobody else, in its own process with its own memory. No shared batch with other customers, so no shared queue and no shared KV cache either.

Logging
Prompts are not logged

The inference layer logs metadata: timestamp, model, number of tokens, latency, status code. The content of prompts and responses is not written anywhere, not even in debug logs. This is the technical side of what is in the addendum.

Metering
Power per card and per server

The consumption of the GPUs and the server is measured, attributed to tokens for shared and to your cards for dedicated, and topped up with the data centre's share via the PUE of our hall. On request in a monthly report.

Maintenance
Updates in a maintenance window

Model versions, drivers and the inference layer are updated in a maintenance window we announce in advance. A new model version gets a new name alongside the old one; you decide when your client switches.

MCP
Tools in your own environment

The model can call MCP tools running in your own environment, for example against the API of a FortiGate or UniFi Dream Machine Pro. Read-only by default with its own revocable account; a change only after a person approves it. The traffic can run entirely over a cross-connect or private VRF.

Access

Four routes to the model.

The endpoint is in BIT-2C in Ede. How your traffic gets there you choose per application; the four routes can exist side by side.

  1. 01Internet
    TLS over the public internet

    The simplest: an HTTPS endpoint with a key. Your traffic goes over the internet encrypted, but it does go over the internet. Suitable for getting started and for applications without sensitive content.

  2. 02WireGuard
    A tunnel to your own environment

    A WireGuard tunnel between your environment and our network, so that the endpoint is not publicly reachable and only your peers can access it. The traffic still goes over the internet, but the endpoint is no longer exposed.

  3. 03Cross-connect
    Your servers in the same data centre

    Your servers in colocation with us at BIT, with a cross-connect to the inference. Model, vector database and application in the same building; the traffic does not go over the internet and there are no egress costs.

  4. 04Private VRF
    Over our EVPN network, from your other locations

    A private VRF on our EVPN-VXLAN fabric, routed to the endpoint. That way you reach the model from your rack in another BIT data centre or at one of our core locations, without the traffic leaving our own network.

How we work

From key to production.

  1. Intake and model choice
    Task, data, expected tokens per day and the latency you need. That leads to a model, a quantisation and the choice between shared and dedicated.
  2. Addendum and data processing agreement
    Data processing agreement and the addendum on training and storage, signed before any production data goes to the endpoint.
  3. Setting up access
    TLS key, WireGuard peer, cross-connect or VRF. With a VRF we walk through the routing with you and test from your side.
  4. Production
    Setting limits, making usage visible, and an engineer who calls you when a model version changes.

A trial setup on the shared capacity, with test data, is possible before you sign.

Who it is for

This page is for you if you:

  • Build the integration yourself and want to know what compatible means exactly
  • Are designing a RAG pipeline and want embeddings and the chat model in one environment
  • Must be able to explain what is logged and what is not
  • Want the traffic to run over a VRF or cross-connect and not over the internet
Why Xyphen IT

Built by network people.

  • EVPN-VXLAN fabric and private VRFs are our daily work, not an extra option
  • Our own DDoS filtering at the edge, also for your endpoint if it is publicly reachable
  • Core locations at NorthC and Nikhef, rack space in all BIT data centres in Ede
  • An engineer on the line, including about quantisation and context

With us, inference is a service on the network, not a separate cloud next to it.

A question about the API or the network?

Call or mail an engineer. If you want a proposal, tell us your task and your expected usage; you receive an outline with model, form and access.