Documentation menu
Requirements

Requirements

Everything the ADP KV cache needs at a glance — the GPU server, the ADP card and its SSDs, the card's software, the network of a remote card, and the software from Awide Labs.

At a glance

PartWhat it needsDetails
GPU serverx86_64 with NVIDIA GPUs vLLM supports (compute capability 7.5+); Linux; NVIDIA driver R580+ (CUDA 13.0 images) or R575+ (CUDA 12.9); Docker 25+ with the NVIDIA Container Toolkit, or Kubernetes; /dev/shm for a 6–10 GiB ringGPU server
ADP cardA PCIe Gen5 x8 slot; low-profile HHHL card with tall and short brackets; under 25 W typical, 45 W at most from the slot; 10–52 °C at 250 LFM airflowADP card hardware
Cache SSDsNVMe SSDs in the card’s server, on the card’s NUMA node, dedicated to it; one RAID 0 array; up to 128 TB raw per card, about 90 % usableSSDs
Card softwareThe ADP software stack (driver, firmware, storage service, CLI, and for a remote card the NVMe-oF target) built for the host’s kernel; prerequisite packages; key eviction onADP card software
Storage server (remote card)x86_64; dedicated cores and hugepages for the NVMe-oF target; one RDMA NIC per card on its NUMA node; a separate boot diskStorage server
Network (remote card)NVMe-oF over RDMA (RoCE v2 or InfiniBand), 100 Gb/s+ per card, MTU 9000 end to end, port 4420 reachable from the GPU server onlyNetwork
Software from Awide LabsThe vLLM image with the connector (or the client kit) and the gateway image of the same release, from nexus.awide.ioInstallation

Local or remote card

Local cardRemote card
Card and SSDsin the GPU serverin a storage server
ADP software stackon the GPU serveron the storage server
On the GPU serverthe stack, the gateway, vLLMthe NVMe-oF initiator, the gateway, vLLM
NetworknoneRDMA, 100 Gb/s+
Gateway backendlocal-xdpremote-xdp

Architecture compares the two, and ADP card setup installs either.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs