This page installs the recommended integration, the vLLM KV connector, with Docker Compose on one GPU server. Kubernetes takes the same pieces; the LMCache plugin and the OffloadingSpec backend have pages of their own.
1. Prepare the card
Set up the card as described in ADP card setup and check that pliocli system get_status reports
Aggregated status is OK and key eviction is enabled. Then note the device the gateway will use:
| Placement | Device for the gateway |
|---|---|
| Local card | /dev/pliops0 (card 0), and the log directory /var/log/pliops |
| Remote card | the NVMe-oF namespace connected on the GPU server, for example /dev/nvme1n1 (nvme list, model XKV_LIGHTNING_AI) |
2. Get the software
The images and client kits are published in the Awide Labs registry. The account is issued on request — contact us to get one.
docker login nexus.awide.io
| Artifact | Coordinates |
|---|---|
| vLLM 0.29.0 with the connector | nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.29.0-6.1.1 |
| vLLM 0.28.0 with the connector | nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1 |
| Storage gateway (remote card) | nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1 |
| Client kit for vLLM 0.29.0 | https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.29.0-6.1.1.tar.gz |
| Client kit for vLLM 0.28.0 | https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.28.0-6.1.1.tar.gz |
Pick the vLLM version your model is recommended on (Supported models). Use the gateway and the vLLM
image (or kit) of the same release: at the handshake the gateway refuses an SDK of another version
(VERSION_MISMATCH in its log), and the connector then stays off. The gateway with the local-card backend is delivered
together with the ADP software stack. The product image is the
upstream vllm/vllm-openai image of that version with the SDK and the connector added; nothing else in it is changed,
and it starts the OpenAI-compatible server like the upstream image does.
docker pull nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
docker pull nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
Building on your own vLLM image. If you run a vLLM image of your own — a CUDA 12.9 variant, a build with your patches — use the client kit instead of the product image. It compiles the SDK against the Python, PyTorch and CUDA of that image and installs the connector on top, without changing the image’s packages:
curl -fu <user> -O https://nexus.awide.io/repository/awide-ai-bundles/awideai/akv-kit-cuda-v0.28.0-6.1.1.tar.gz
tar xzf akv-kit-cuda-v0.28.0-6.1.1.tar.gz && cd akv-kit-cuda-v0.28.0-6.1.1
./build.sh --base-image vllm/vllm-openai:v0.28.0-cu129 --tag my-registry/vllm-openai-awideai:6.1.1
3. Write the compose file
The gateway and vLLM run as two containers. The gateway creates the shared-memory ring as a file in /dev/shm, so
both must use the host’s /dev/shm: ipc: host on both, and nothing mounted over /dev/shm in either. vLLM
reaches the gateway’s control port by its service name (XKV_GATEWAY_HOSTNAME).
The file below is for a remote card and GLM-5.3 FP8 on eight GPUs; the values that depend on the model are marked, and Supported models has them for the other models.
services:
gateway:
image: nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
ipc: host # the ring is a file in the host's /dev/shm
cap_add: [SYS_ADMIN, IPC_LOCK] # SYS_ADMIN: NVMe pass-through commands to the card
ulimits:
memlock: -1
devices:
- /dev/nvme1n1:/dev/nvme1n1 # the connected ADP namespace (nvme list)
command:
- --storage-backend=remote-xdp
- --storage-devices=/dev/nvme1n1
- --fragment-size=131072 # the NVMe-oF transfer size of the card
- --port=5557
- --shared-memory-size=10GiB
- --thread-pool-size=64
- --num-dbs=128
- --start-clean=0 # keep the cache across restarts
restart: unless-stopped
vllm:
image: nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
ipc: host
cap_add: [IPC_LOCK]
ulimits:
memlock: -1
depends_on: [gateway]
ports:
- "8000:8000"
deploy:
resources:
reservations:
devices:
- { driver: nvidia, count: all, capabilities: [gpu] }
environment:
PYTHONHASHSEED: "0" # required: otherwise nothing is found after a restart
XKV_GATEWAY_HOSTNAME: gateway # where the SDK finds the gateway
# VLLM_USE_V2_MODEL_RUNNER: "0" # required on vLLM 0.29
volumes:
- /models:/models:ro
command:
- --model=/models/GLM-5.3-FP8
- --tensor-parallel-size=8
- --enable-prefix-caching
- --enable-prompt-tokens-details
- --kv-cache-dtype=fp8 # model-specific
- --max-num-batched-tokens=8192 # model-specific
- '--kv-transfer-config={"kv_connector":"FusIOnXConnector","kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both","kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,"slot_alignment_bytes":4096}}'
restart: unless-stopped
Keep the JSON of --kv-transfer-config (and of --speculative-config, if you use MTP) in single quotes as above, and
add the model’s own options — tool-call and reasoning parsers, --max-model-len — as usual.
For a local card, only the gateway changes: the gateway image with the local backend, the card’s device and log directory instead of the NVMe namespace, and the local backend’s default fragment size.
gateway:
image: <local-card gateway image, release 6.1.1> # delivered with the ADP software stack
ipc: host
cap_add: [IPC_LOCK]
ulimits:
memlock: -1
devices:
- /dev/pliops0:/dev/pliops0
volumes:
- /var/log/pliops:/var/log/pliops
command:
- --storage-backend=local-xdp # fragment size defaults to 261120 here
- --port=5557
- --shared-memory-size=10GiB
- --thread-pool-size=64
- --num-dbs=128
- --start-clean=0
restart: unless-stopped
On a local card the connector needs no slot_alignment_bytes. Every gateway setting can also come from a YAML file
(--config /etc/akv/gateway.yaml); flags given on the command line win. All of them:
Configuration.
4. Start
docker compose up -d
docker compose logs -f gateway vllm
curl -f http://127.0.0.1:8000/health
The first start of a large model takes a while: weights load, CUDA graphs are captured. When the connector has registered, the vLLM log shows which caches it exchanges and how:
XKV: <N> cache tensors registered for exchange
KV cache registry: <N> cache(s) in <G> group(s); writers: layer_hook=... ; readers: ...
A model whose caches the connector cannot account for does not start at all, and the log names the cache — the connector never runs half-configured.
5. Verify a restore
A restore is proved by a hit that can only come from the card: send a long prompt, restart only vLLM (the gateway and the card keep the cache), and send it again.
MODEL=/models/GLM-5.3-FP8
PROMPT=$(python3 -c "print(' '.join(f'Line {i}: the ADP card keeps this prompt.' for i in range(3000)))")
ask() {
curl -s http://127.0.0.1:8000/v1/completions -H 'Content-Type: application/json' \
-d "$(jq -n --arg m "$MODEL" --arg p "$PROMPT" '{model: $m, prompt: $p, max_tokens: 8}')" \
| jq '.usage | {prompt_tokens, cached: .prompt_tokens_details.cached_tokens}'
}
ask # cold: cached is 0, the prompt is computed and written to the card
docker compose restart vllm # GPU memory is empty now, the card is not
until curl -sf http://127.0.0.1:8000/health; do sleep 10; done
time ask # warm: cached is close to prompt_tokens, and the answer comes much faster
cached_tokens is rounded down to whole chunks: with chunk_size 128 a prompt of 30,100 tokens can hit at most
30,080. A prompt shorter than one chunk is never stored. Over time the vLLM log reports the hit rate of the card, and
/metrics exposes the same counters:
External prefix cache hit rate: 99.6%
vllm:external_prefix_cache_queries_total vllm:external_prefix_cache_hits_total
If cached_tokens stays at 0 after the restart, see Operations & troubleshooting.
Stopping and upgrading
docker compose down keeps the cache on the card: with --start-clean=0 the next start finds it. An upgrade of the
images keeps it too, unless the release notes say the stored layout changed — then wipe the card as described in
ADP card setup. Changing the model, chunk_size or the tensor-parallel size starts
a new cache in any case (Configuration).
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.