Documentation menu
Get started

Kubernetes

What a vLLM pod needs to use the ADP card: the gateway as a sidecar, the node's shared memory, access to the card, and a restart policy that brings both back together.

On Kubernetes the gateway runs as a second container in the vLLM pod. Everything from Installation carries over; what changes is how the two containers share memory and how they are restarted. Any operator or chart that produces the pod below works the same way.

What the pod needs

RequirementHow
The gateway next to vLLMA second container in the same pod: one gateway per vLLM instance.
One /dev/shm for bothhostIPC: true on the pod and no emptyDir (or any other volume) mounted at /dev/shm in either container. A private /dev/shm shows up in vLLM as Failed to open KVDatabase while the gateway reports the memory allocated.
A unique gateway port per nodeThe ring is a file in the node’s /dev/shm named after the gateway port, so two pods on one node need different --port values (and the matching xkv_gateway_port).
Access to the cardThe gateway container runs privileged to reach the card’s device (/dev/pliops0 for a local card, the connected NVMe-oF namespace for a remote one) and send it pass-through commands. A remote card is connected on the node, outside the pod (ADP card setup).
The control port on loopbackContainers of one pod share the network namespace: bind the gateway to 127.0.0.1 (--bind-address=127.0.0.1). The SDK connects to localhost by default.
A fixed hash seedPYTHONHASHSEED=0 in the vLLM container.
Pulling from the registryAn image pull secret for nexus.awide.io.
SchedulingThe pod must land on a node that has the card (local) or the connected namespace (remote): a node selector or affinity on a label you set on those nodes.

Example

kubectl create secret docker-registry awide-registry \
  --docker-server=nexus.awide.io --docker-username=<user> --docker-password=<token>
apiVersion: apps/v1
kind: Deployment
metadata:
  name: glm53-adp
spec:
  replicas: 1
  selector:
    matchLabels: { app: glm53-adp }
  template:
    metadata:
      labels: { app: glm53-adp }
    spec:
      hostIPC: true                              # the ring lives in the node's /dev/shm
      nodeSelector:
        example.com/adp-card: "true"             # a label you put on the nodes with the card
      imagePullSecrets:
        - name: awide-registry
      containers:
        - name: gateway
          image: nexus.awide.io/awide-docker-hosted/akv-gateway:6.1.1
          args:
            - --storage-backend=remote-xdp
            - --storage-devices=/dev/nvme1n1
            - --fragment-size=131072
            - --bind-address=127.0.0.1
            - --port=5557
            - --shared-memory-size=10GiB
            - --thread-pool-size=64
            - --num-dbs=128
            - --start-clean=0
          securityContext:
            privileged: true                     # the card's device and NVMe pass-through
        - name: vllm
          image: nexus.awide.io/awide-docker-hosted/awideai-cuda-vllm-openai:v0.28.0-6.1.1
          args:
            - --model=/models/GLM-5.3-FP8
            - --tensor-parallel-size=8
            - --enable-prefix-caching
            - --enable-prompt-tokens-details
            - --kv-cache-dtype=fp8
            - --max-num-batched-tokens=8192
            - '--kv-transfer-config={"kv_connector":"FusIOnXConnector","kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both","kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,"slot_alignment_bytes":4096}}'
          env:
            - { name: PYTHONHASHSEED, value: "0" }
          ports:
            - containerPort: 8000
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            periodSeconds: 10
          resources:
            limits:
              nvidia.com/gpu: 8
          volumeMounts:
            - { name: models, mountPath: /models, readOnly: true }
      volumes:
        - name: models
          hostPath: { path: /models }

For a local card, the gateway container takes the local-card gateway image delivered with the ADP software stack, --storage-backend=local-xdp instead of the three NVMe-oF options, and a hostPath volume for /var/log/pliops. The gateway and vLLM images always come from the same release.

Check the shared memory once after the first start:

kubectl exec deploy/glm53-adp -c gateway -- sh -c 'echo ok > /dev/shm/adp-shm-check'
kubectl exec deploy/glm53-adp -c vllm -- cat /dev/shm/adp-shm-check     # must print "ok"
kubectl exec deploy/glm53-adp -c gateway -- rm /dev/shm/adp-shm-check

Restarts

The gateway and vLLM form one unit: vLLM opens its connection to the gateway when it starts. Restart them together.

  • When the gateway container restarts, recreate the pod (kubectl delete pod …), so that vLLM connects to the new gateway. Until then vLLM serves from GPU memory alone and logs External KV cache turned off on this rank.
  • When the vLLM container restarts on its own, with the gateway running, nothing else is needed — that is exactly the restore path of Installation.

Recreating the pod on a restart of the gateway container is easy to automate: alert on the restart count of the gateway container or on the log line above.

Security

The gateway’s control port answers lookups and deletes without authentication, and /metrics and the KV events of vLLM show cache behaviour. Keep the gateway on loopback as above, and do not expose /metrics, KV events or the gateway port outside the pod. For multi-tenant serving, see Tenant isolation.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs