Documentation menu
Requirements

GPU server

What the server that runs vLLM needs for the ADP KV cache: GPUs, operating system, driver and containers, host resources, and the software that comes from Awide Labs.

Hardware

Requirement
Serverx86_64 with NVIDIA GPUs that vLLM supports.
GPUsThe upstream vLLM images are built for compute capability 7.5 and newer: Turing, Ampere, Ada, Hopper, Blackwell. Some models need more than that from vLLM itself — DeepSeek Sparse Attention models run only on Hopper and Blackwell.
GPU memoryWhat the model and its KV cache need without the connector, plus a small headroom: the connector’s buffers are allocated after vLLM sizes its KV cache (resources).
ADP cardLocal placement: a free PCIe Gen5 x8 slot and the SSDs behind it, see ADP card hardware. Remote placement: nothing of the card here, only the network card below.
Network (remote card)An RDMA-capable NIC (RoCE v2 or InfiniBand), 100 Gb/s or faster, jumbo frames (MTU 9000), see network.

Operating system and software

Requirement
OSLinux x86_64
NVIDIA driverOne that supports the CUDA of the vLLM image: R580 or newer for the CUDA 13.0 images (the default vllm/vllm-openai tags), R575 or newer for the CUDA 12.9 images (-cu129 tags).
ContainersDocker Engine 25 or newer with Compose v2 and the NVIDIA Container Toolkit (nvidia runtime), or Kubernetes with the NVIDIA device plugin.
Local cardThe ADP software stack on this server (ADP card software).
Remote cardnvme-cli; the kernel modules nvme_rdma and rdma_ucm; the RDMA user-space stack of your NIC; NVMe native multipath off (kernel parameter nvme_core.multipath=N).

Everything else — CUDA, PyTorch, Python, vLLM, the connector and the SDK — comes inside the container image. The connector is installed into the vLLM image without changing any of its packages.

Resources

ResourceWhat to plan for
/dev/shmAt least the ring of every gateway on the host (shared_memory_size, 6 GiB by default, 10 GiB for 8-GPU servers), plus what vLLM uses itself. The gateway and vLLM must share the host’s /dev/shm.
Host RAMThe ring above plus vLLM’s own needs. The connector keeps no cache in host memory.
CPUA few cores for the gateway: its worker threads (thread_pool_size, 4 by default, 64 on 8-GPU servers) each run one device command at a time.
DiskAbout 50 GB for the images, plus the model weights.
GPU memory headroomNVLink fan-out, when enabled, takes a 256 MiB staging buffer and an NCCL communicator per GPU outside gpu_memory_utilization; lower --gpu-memory-utilization by about 0.005 on a 141 GB GPU for the defaults. Long prompts on large models may need a little more headroom than without the connector.

Software from Awide Labs

ArtifactWhat it is
Product imagevLLM with the SDK and the connector installed, for vLLM 0.28.0 and 0.29.0
Gateway imageThe storage gateway for a remote card (and a directory backend for trials). The gateway with the local-card backend comes with the ADP software stack.
Client kitSDK and connector sources with a Dockerfile, to build on your own vLLM image

The images and kits are published in the Awide Labs registry nexus.awide.io; the account is issued on request. The gateway and the vLLM image (or kit) must come from the same release: the gateway refuses an SDK of another version. Installation has the exact names.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs