At a glance
| Part | What it needs | Details |
|---|---|---|
| GPU server | x86_64 with NVIDIA GPUs vLLM supports (compute capability 7.5+); Linux; NVIDIA driver R580+ (CUDA 13.0 images) or R575+ (CUDA 12.9); Docker 25+ with the NVIDIA Container Toolkit, or Kubernetes; /dev/shm for a 6–10 GiB ring | GPU server |
| ADP card | A PCIe Gen5 x8 slot; low-profile HHHL card with tall and short brackets; under 25 W typical, 45 W at most from the slot; 10–52 °C at 250 LFM airflow | ADP card hardware |
| Cache SSDs | NVMe SSDs in the card’s server, on the card’s NUMA node, dedicated to it; one RAID 0 array; up to 128 TB raw per card, about 90 % usable | SSDs |
| Card software | The ADP software stack (driver, firmware, storage service, CLI, and for a remote card the NVMe-oF target) built for the host’s kernel; prerequisite packages; key eviction on | ADP card software |
| Storage server (remote card) | x86_64; dedicated cores and hugepages for the NVMe-oF target; one RDMA NIC per card on its NUMA node; a separate boot disk | Storage server |
| Network (remote card) | NVMe-oF over RDMA (RoCE v2 or InfiniBand), 100 Gb/s+ per card, MTU 9000 end to end, port 4420 reachable from the GPU server only | Network |
| Software from Awide Labs | The vLLM image with the connector (or the client kit) and the gateway image of the same release, from nexus.awide.io | Installation |
Local or remote card
| Local card | Remote card | |
|---|---|---|
| Card and SSDs | in the GPU server | in a storage server |
| ADP software stack | on the GPU server | on the storage server |
| On the GPU server | the stack, the gateway, vLLM | the NVMe-oF initiator, the gateway, vLLM |
| Network | none | RDMA, 100 Gb/s+ |
| Gateway backend | local-xdp | remote-xdp |
Architecture compares the two, and ADP card setup installs either.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.