Documentation menu
Requirements

ADP card

What installing the ADP card takes and what it runs on: the slot, power and cooling, the SSDs behind it, placement, the storage server and network of a remote card — and the ADP software stack, the operating system, packages, services and settings it needs.

Hardware

The card

Requirement
SlotPCIe Gen5, 8 lanes (the card is a PCIe Gen5 x8 device).
Form factorLow-profile, half-height half-length (HHHL, 6.6” × 2.536”); with tall and short brackets for full-height and low-profile slots.
PowerUnder 25 W typical, 45 W at most, +12 V through the PCIe slot — no auxiliary power cable.
CoolingOperating temperature 10–52 °C at 250 LFM of airflow across the card; storage 5–35 °C, below 90 % humidity, non-condensing.
Power-loss protectionOn-board supercapacitors; no battery and nothing to install.
ServersAny standard x86_64 server with a free slot: Dell, HPE, Lenovo, Supermicro, Quanta, Wiwynn, Inspur, Sugon, Fujitsu, Hitachi, Tyan, MiTAC, Intel, Cisco, AIC.

GPU servers often run their fans by GPU temperature; check that the slot you choose gets the airflow above when the GPUs are idle too.

The SSDs behind the card

The card stores the cache on NVMe SSDs in the same server. At installation the SSDs are assigned to the card and combined into one array behind it.

Requirement
InterfaceNVMe (PCIe Gen3, Gen4 or Gen5). The card also drives SAS and SATA SSDs; for the KV cache use NVMe.
Type and vendorTLC or QLC flash from any vendor (Samsung, WD, Micron, Intel, Kioxia, Hynix, Seagate, …).
ArrayRAID 0 across the SSDs of one card. The content is a cache that can always be recomputed, so there is no need to pay for redundancy.
CapacityUp to 128 TB of raw SSD capacity per card. The usable cache is about 90 % of the raw capacity: two 3.84 TB SSDs give about 7 TB.
DedicatedThe SSDs belong to the card: no file system, no partitions, not the boot disk. Keep the operating system on a disk of its own.
PlacementOn the same NUMA node (CPU socket) as the card.

How much capacity. Plan from the prefixes you want to keep warm, not from the request rate: the KV size per token of your model times the tokens of all long prompts, conversations and agent contexts that should come back without a prefill. The usage report of the card (pliocli system get_disk_usage) shows the stored objects and their average size once a model has run for a while. When the card is full it evicts on its own and the hit rate is what the capacity allows (capacity and key eviction).

Where the card goes

PlacementHardware
Local — in the GPU serverA free PCIe Gen5 x8 slot in the GPU server and room for the SSDs there. No network.
Remote — in a storage serverA storage server with the card and its SSDs (below), and an RDMA network between it and the GPU server. The GPU server needs only the network card.

A remote card keeps the GPU servers free of storage and lets the cache sit on hardware sized for it; a local card is the simplest setup, with nothing between the gateway and the card.

Several cards in one server

Put one card per NUMA node: in a two-socket server, a card on each socket, each with its own SSDs on that socket. Every card has its own storage service and, in a remote setup, its own NVMe-oF target and network address. Give each gateway a card of its own; several vLLM instances on one GPU server can each use a separate card.

Storage server for a remote card

Requirement
CPUx86_64. The NVMe-oF target service polls its CPU cores continuously, so give every card a set of dedicated cores on its own NUMA node — they show as 100 % busy, which is normal.
MemoryEnough for the operating system plus the hugepages the installer reserves for the target service.
NetworkAn RDMA-capable NIC per card, on the card’s NUMA node (below).
DisksThe SSDs of the cards, and a separate boot disk.
Remote managementIPMI or another out-of-band console is strongly recommended: the target runs on bare metal.

Network for a remote card

Requirement
TransportNVMe over Fabrics on RDMA: RoCE v2 over Ethernet, or InfiniBand.
NICRDMA-capable, on the storage server and on the GPU server; on the storage server on the same NUMA node as its card.
Speed100 Gb/s or faster per card. The restore speed of a long prompt is capped by the slower of the link and the SSDs.
MTUJumbo frames, 9000, end to end — NICs and switch ports alike.
AddressingOne IP address and port (4420) per card’s target; the GPU server must reach it, tenants must not.

A slower link works, just with slower restores; for RoCE, configure the switches the way your NIC vendor recommends for RDMA traffic.

Software

The ADP software stack

Everything that runs the card comes in one package from Awide Labs, the ADP software stack. It is installed on the host of the card — the GPU server for a local card, the storage server for a remote one.

ComponentWhat it isOn the host as
DriverThe kernel module of the card. Built for a specific kernel version.kernel module
FirmwareThe card’s firmware, delivered with the stack.reported by pliocli system get_status
Storage serviceKeeps the key-value databases on the card; one instance per card. Starts at boot.pliostore@<N> (systemd), config /etc/pliops/<N>/pliostore.ini
Command-line toolStatus, usage, the SSD array, maintenance.pliocli
NVMe-oF target service (remote card only)Exports the card’s key-value namespace over RDMA, one instance per card.lightning-spdk-target-<N> (systemd), port 4420
InstallersInstall the above and build the SSD array; the target installer sets up the NVMe-oF target.xdp-installer, lightning_ai_spdk_target.py

On the GPU server, the connector and the gateway come as container images (Installation). With a local card the gateway container talks to the storage service on the same host; with a remote card nothing of the stack runs on the GPU server — only the kernel’s NVMe-oF initiator.

Release 6.1.1 of the connector is verified with ADP software stack 2.4.5.

Operating system and kernel

HostRequirement
Host of the cardLinux x86_64. The driver is built for the exact kernel the host runs: tell Awide Labs your distribution and kernel version and you get the stack built for it. A kernel update needs a stack built for the new kernel — pin the kernel on these hosts.
GPU server, remote cardLinux x86_64 with the NVMe-oF RDMA initiator in the kernel (nvme_rdma) — any current distribution kernel has it.

Packages to install first

On the host of the card the installers need these packages (names as in RHEL-family distributions; the Debian family has equivalents):

PackageVersionWhy
nvme-cli, pciutils, libpciaccessanyFinding and managing the card and its SSDs.
libaio, numactl, numactl-develanyI/O and NUMA placement of the storage service.
ledmon0.96 or newerDrive LEDs of the SSD array.
fuse3, fuse3-devel, libuuid-devel, ncurses-develanyRequired by the stack’s services and tools.
python3, python3-pip3.8 or newerThe installers; pip modules meson and pyelftools for the target installer.
libibverbs, libibverbs-devel, librdmacm-develanyRDMA for the NVMe-oF target (remote card).
tunedanyThe throughput-performance profile of a storage server.

The target installer also installs the dependencies of the NVMe-oF target with RDMA support itself.

On the GPU server with a remote card:

Package or settingWhy
nvme-clinvme discover, nvme connect, nvme list.
kernel modules nvme_rdma, rdma_ucmThe NVMe-oF initiator over RDMA.
nvme_core.multipath=N (kernel parameter)The gateway sends NVMe pass-through commands to the namespace device; native NVMe multipath must be off.
the RDMA user-space stack of your NIC (rdma-core or the vendor’s driver package)RDMA on the NIC.

Installing and checking it

ADP card setup installs the stack, checks the card and its services, and sets the two storage service keys the KV cache depends on — key eviction on, key-value mode. Every key is listed in Configuration → ADP storage service.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs