Documentation menu
Get started

ADP card setup

Installing and configuring the card itself: the software on its host, the storage service settings that matter, capacity and key eviction, wiping the cache, health checks, and connecting a remote card over NVMe-oF.

What runs on the card’s host

The card is managed by the ADP software stack on the server it sits in — the GPU server for a local card, a storage server for a remote one: a driver, one storage service per card (pliostore@<N>), the pliocli tool and, for a remote card, one NVMe-oF target service per card (lightning-spdk-target-<N>). What each part is and what the host needs for it: ADP card requirements.

The software stack in both placements
what the gateway talks to
LOCAL CARD GPU SERVER Storage gateway local backend ADP storage service key-value databases · one per card ADP driver ADP card · PCIe slot NVMe SSDs behind it, one array objects up to 256 KiB REMOTE CARD GPU SERVER Storage gateway · remote backend NVMe-oF initiator · /dev/nvmeXnY RDMA fabric NVMe-oF, port 4420 STORAGE SERVER NVMe-oF target service · 128 KiB I/O ADP storage service + driver ADP card · NVMe SSDs behind it Nothing in vLLM changes between the two: only the gateway backend and its device.

Install the software stack

The ADP software stack is delivered by Awide Labs as an installer built for the kernel of the card’s host. What the host needs before you start — slot, SSDs, packages, kernel — is in ADP card requirements.

  1. Install the card in a PCIe Gen5 x8 slot and the NVMe SSDs that will hold the cache in the same server, on the card’s NUMA node.

  2. Install the prerequisite packages.

  3. Run the installer with the SSDs of the card listed explicitly. It installs the driver, the storage service and the command-line tool, puts the SSDs into one RAID 0 array behind the card and reserves hugepages:

    sudo ./xdp-installer install --level=0 --resources="/dev/nvme1n1 /dev/nvme2n1" --vhc=0 \
      --enable-set-hugepages --hugepage-distribution-percentage=50 --assume-yes
  4. Remote card only: install the NVMe-oF target on the address the GPU server will connect to, with the 128 KiB I/O size the gateway’s fragments are sized for, and start it:

    ./lightning_ai_spdk_target.py install --targets=192.0.2.10:4420 --start-service --io-size 131072
  5. On a storage server, set the CPU and I/O tuning profile:

    sudo tuned-adm profile throughput-performance
  6. Check that the card is up (next section), then configure the storage service.

Check the card

systemctl is-active pliostore@0          # the storage service of card 0: "active"
pliocli system get_xdp_list              # the cards on this host, their device paths and mode
pliocli system get_status                # every component, and the aggregated status
pliocli system get_disk_usage -x 0       # physical usage, valid objects, databases
pliocli raid get_raid_info -x 0          # the SSD array behind card 0
systemctl is-active lightning-spdk-target-0   # remote card only: the NVMe-oF target of card 0

A healthy card reports every component OK:

Components status:
	Driver (OK): Driver is working properly
	Service (OK): Pliostore service is working properly and system is active
	RAID (OK): RAID Array is functioning properly, raidLevel: RAID0, protection: No, performance: Normal
	CapacityUsage (OK): Capacity usage is below defined thresholds
	Firmware (OK): Firmware is working properly
	Temperature (OK): Temperature is at NORMAL level
	SuperCapacitors (OK): Super Capacitors are working properly
	OnboardFlash (OK): Onboard flash component is working properly
Aggregated status is OK

The usage report shows how full the card is and how many objects and databases it holds; the number of databases is the gateway’s num_dbs (128 by default) once a gateway has connected:

Physical usage [%]       : 12.40
Physical capacity        : 863.10 GB / 6.96 TB
Number of valid objects  : 10051238 / 1680330742
Number of databases      : 128 / 4096 (0 deleting)

Configure the storage service

The storage service of card N reads /etc/pliops/<N>/pliostore.ini, which the installer writes. Check two keys on every card that serves the KV cache — the full list is in Configuration → ADP storage service:

[db_settings]
key_eviction = true
product_type = 0                        # 0 = KV, 1 = Block

Apply a change with the model server and the gateway stopped:

systemctl restart pliostore@0
journalctl -u pliostore@0 -o cat | grep -i "key eviction"    # "Key eviction Enabled"
systemctl restart lightning-spdk-target-0                     # remote card: restart the target after the service

A remote card then has to be reconnected on the GPU server.

Capacity and key eviction

The storage service reports usage in three zones, visible as CapacityUsage in pliocli system get_status:

Physical usageStatusEffect
below the soft threshold (90 % by default)OK—
between the soft and the hard threshold (100 %)warningnone on writes
above the maximumCRITICALevery incoming write fails until space is freed; the service logs event 3002

With key_eviction = true the card starts reclaiming space before it reaches the maximum. The point where eviction starts is built into the storage service, not a setting in the INI file, and the soft and hard thresholds (pliocli system get_usage_soft_threshold, get_usage_hard_threshold) only change the reported status — do not move them expecting eviction to start earlier.

Watch the physical usage (pliocli system get_disk_usage) like the fill level of any cache, and alert well before the maximum: a card that fills faster than it reclaims shows up as a falling hit rate, not as an error in vLLM.

Wiping the card

A wipe empties the cache: after a change of the stored layout (settings that require a clean cache) it is required, otherwise it only costs a refill. The gateway does it at start with --start-clean=1, by deleting its databases — which the storage service refuses while key eviction is on (Key eviction is ENABLED, delete DB not supported). So a wipe takes this order:

  1. Stop vLLM and the gateway.
  2. Set key_eviction = false and restart the storage service (and the target, then reconnect — remote card).
  3. Start the gateway once with --start-clean=1 and wait until its log reports the databases deleted; pliocli system get_disk_usage -x 0 then shows 0 objects.
  4. Set the gateway back to --start-clean=0.
  5. Set key_eviction = true and restart the storage service (and the target, then reconnect).
  6. Start the gateway and vLLM.

WARNING

With key eviction enabled, keep start_clean at 0 everywhere the gateway is started — compose files, Kubernetes manifests, restart scripts. A gateway that starts with start_clean=1 on every restart would otherwise try to wipe the card each time.

Remote card: connect the GPU server

The NVMe-oF target service on the storage server exports one key-value namespace per card over RDMA, port 4420. On the GPU server, connect it with the kernel initiator:

modprobe nvme_rdma
nvme discover -t rdma -a 192.0.2.10 -s 4420                 # lists the subsystem NQN of the card
nvme connect  -t rdma -a 192.0.2.10 -s 4420 -n <subsystem NQN>
nvme list          # the ADP namespace, model XKV_LIGHTNING_AI
nvme list-subsys   # ... rdma traddr=192.0.2.10,trsvcid=4420 live

The namespace is a key-value namespace, not a block device: nvme list shows its size as 0 B and its SMART log is empty — both are normal. Do not partition or format it. Its device path (/dev/nvme1n1, for example) is what the gateway gets as --storage-devices, and the gateway container needs that device and SYS_ADMIN for the NVMe pass-through commands (Installation).

Make the connection persistent the way your distribution does it (for example /etc/nvme/discovery.conf and the nvmf-autoconnect service of nvme-cli), and check the device name after every reconnect: it can change.

After a restart of the target

A restart of the storage service or of the target on the storage server can change the namespace identifiers, and the connected device stops answering. On the GPU server:

nvme disconnect -n <subsystem NQN>
nvme connect -t rdma -a 192.0.2.10 -s 4420 -n <subsystem NQN>
nvme list          # the device path may differ from before

Then restart the gateway and vLLM together, so that both open the device again.

Transfer size and fragments

PlacementLimitGateway fragment_size
Local cardan object of at most 256 KiB261120 (255 KiB; the default of the local backend)
Remote cardone NVMe command of at most MDTS — 128 KiB with the target’s --io-size 131072 — and a multiple of 4 KiB131072

The gateway logs the device limit at start (# Max I/O Size) and refuses a fragment size that does not fit it. Models whose KV page is not a multiple of 4 KiB (GLM-5.x) also need slot_alignment_bytes: 4096 in the connector settings on a remote card. Changing fragment_size requires a wipe.

Durability

The card holds a cache, and the system is built to lose nothing that matters if it loses some of it: a write the card has acknowledged can still be missing after an abrupt restart of the storage service or the target, and the connector then reads that chunk as a miss and vLLM recomputes it. The card itself protects its data against power loss with on-board supercapacitors.

Reading the gateway log

NVMe media error type: 2 code: 135 is how the card answers a lookup of a key it does not hold — an ordinary cache miss, which the gateway logs at ERROR level. Thousands of them during a run are normal. The same code is also returned for other refusals, a full card among them, so judge the card by pliocli system get_status and get_disk_usage, not by the count of these lines.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs