How it plugs in
The connector implements vLLM’s KV connector API (V1), including the interface for hybrid models with several KV cache
groups. vLLM loads it by name from --kv-transfer-config; no fork or patch of the engine is involved, and one code
base runs on vLLM 0.26, 0.28 and 0.29 by detecting features rather than versions.
It works on both sides of the engine:
- Scheduler side — for a new request it asks the gateway which chunks of the prompt are stored, tells vLLM how many tokens can be loaded, and after allocation decides which blocks to read and which new ones to write.
- Worker side (every tensor-parallel rank) — it registers the KV caches the engine allocated, issues the reads at the start of the forward pass, fences each layer on its own pages, and writes new pages after each layer. A registry decides at start who writes and who reads every cache the engine allocated; a model with a cache nobody would write is refused at start instead of silently caching nothing.
Architecture follows a request through both.
Enable it
vllm serve <model> \
--enable-prefix-caching \
--enable-prompt-tokens-details \
--kv-transfer-config '{"kv_connector":"FusIOnXConnector",
"kv_connector_module_path":"xkv_vllm_connector.fusionx_connector",
"kv_role":"kv_both",
"kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128}}'
with, in the environment of the vLLM process:
export PYTHONHASHSEED=0 # required
export XKV_GATEWAY_HOSTNAME=localhost # where the gateway runs; localhost is the default
export VLLM_USE_V2_MODEL_RUNNER=0 # vLLM 0.29 only
| Setting | Why |
|---|---|
PYTHONHASHSEED=0 | A fixed hash seed keeps vLLM’s block hashes — and the keys on the card — the same across restarts. |
--enable-prefix-caching | The connector serves prefix-cache hits. Hybrid models need it explicitly: without it vLLM turns the Mamba cache off and the connector refuses to start. |
--enable-prompt-tokens-details | Optional: reports cached_tokens in API responses, which is how you see a restore. |
VLLM_USE_V2_MODEL_RUNNER=0 | vLLM 0.29 only: the connector works with the V1 model runner. |
The names FusIOnXConnector, xkv_vllm_connector and xkv_gateway_port are the connector’s interface and stay as
they are across releases. Model-specific options — --kv-cache-dtype, --max-num-batched-tokens, --watermark,
chunk_size — are in Supported models.
Settings
All keys go into kv_connector_extra_config; every key with its default is in
Configuration → Connector. Most deployments set only three:
xkv_gateway_port— the gateway’s control port,5557by default;chunk_size— tokens per stored object, per model (Supported models);slot_alignment_bytes: 4096— on a remote card, for models whose KV page is not a multiple of 4 KiB (GLM-5.x).
Per-request control
A request can be kept out of the external cache entirely — no lookup, no restore, no write — with
{ "kv_transfer_params": { "xkv_store": false } }
in the request body. A front end sets it for request-scoped traffic whose prefix will never be asked for again: nothing
derived from the prompt reaches the card. Clients must not be able to set kv_transfer_params themselves.
Tenant isolation
The connector keys the card only by vLLM’s block hashes, which carry each request’s cache_salt: when the API front
end salts requests per tenant, a tenant’s scope holds on the card exactly as in GPU memory, with no connector setting.
Two optional switches make it stricter — require_cache_salt keeps unsalted requests out of the card, key_namespace
separates deployments that share one. The front end’s rules, a working front end, the switches and a test:
Tenant isolation.
NVLink fan-out
For MLA models such as GLM-5.x every tensor-parallel rank needs the same latent page. With fan-out, each rank of an NVLink group copies only its share of a layer from host memory, and the group exchanges the rest over NVLink with NCCL right before the layer computes. Host-to-GPU traffic of a restore drops from one copy per rank to one per group.
"kv_connector_extra_config": { "nvlink_fanout": true }
The groups (auto takes the GPUs NVML reports as NVLink peers), the staging buffer and the smallest layer worth
fanning out are set by the nvlink_fanout_* keys in Configuration.
It needs NVIDIA GPUs with NVLink peer access, tensor parallelism above 1 and an MLA model whose latent page is the same
on every rank. Hybrid models (Kimi Linear and K3 included), attention-only models and the block-outermost layout of
DeepSeek-V4 do not qualify; in every such case the fan-out stays off with one warning naming the reason, and the
restore runs exactly as without it. The staging buffer and the
NCCL communicator are allocated after vLLM sized the KV cache, outside gpu_memory_utilization: lower
--gpu-memory-utilization by about 0.005 on a 141 GB GPU for the defaults. Setting NCCL_LAUNCH_ORDER_IMPLICIT=1 in
the vLLM environment is harmless and recommended with NCCL 2.26 or newer.
Logs to expect
At start (rank 0):
XKV: <N> cache tensors registered for exchange
KV cache registry: <N> cache(s) in <G> group(s); writers: ... ; readers: ...
KV cache layout: <G> group(s)
While serving, vLLM’s periodic stats line includes External prefix cache hit rate. If the gateway goes away:
External KV cache turned off on this rank
after which the engine answers zero hits until the gateway and vLLM are restarted together (what happens then).
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.