Documentation menu
Overview

Architecture

The components of the ADP KV cache, how a KV page travels to the card and back, how pages are keyed so the cache survives restarts, and how the card can be placed and shared.

Components

ComponentWhere it runsWhat it does
ADP KV cache connectorinside vLLM, in the scheduler process and in every tensor-parallel workerDecides which blocks of a prompt are already stored, reserves GPU blocks for them, moves pages in and out, and fences those transfers against the model’s forward pass. Implements vLLM’s KV connector API (V1) including the hybrid-model interface.
SDK clientinside vLLM, one per tensor-parallel rankBuilds the storage keys, batches requests, and copies pages between GPU memory and the shared-memory ring on a CUDA stream of its own.
Shared-memory ring/dev/shm of the hostThe transfer window between vLLM and the gateway: pages are copied into ring slots, the gateway reads and writes the slots. It is not a cache — a slot lives for one transfer.
Storage gatewaya separate process (container) on the GPU server, one per vLLM instanceCollects the requests of all ranks of one engine, schedules reads and writes, splits objects larger than the card accepts into fragments, spreads them over the card’s databases, and talks to the card.
ADP carda PCIe slot of the GPU server, or a storage server reached over NVMe-oFStores the objects as key-value pairs on NVMe SSDs behind it.

Control messages between the SDK and the gateway go over ZeroMQ (port 5557 by default); data never does, it moves through the ring.

A restore, step by step

Restore and save
one request · one engine step
vLLM scheduler Connector · worker Storage gateway ADP card LOOKUP EXIST · which chunks of the prompt are stored? chunks 1 … k found reserves GPU blocks for them, schedules the step LOAD load chunks 1 … k GET · one request per layer read fragments objects pages land in the ring ring → GPU on its own CUDA stream; layer i waits for its own pages SAVE PUT · pages after the layer write, asynchronously the step does not wait for the card
  1. Lookup. When a request is scheduled, the connector’s scheduler side asks the gateway which chunks of the prompt are on the card (EXIST). vLLM counts them as computed tokens and reserves GPU blocks for them.
  2. Load. In the forward pass every rank sends one read request per layer (GET). The gateway reads the objects into the ring; the SDK copies them into the reserved GPU blocks on a CUDA stream of its own.
  3. Per-layer fence. Before a layer computes, it waits for its own pages only (wait_for_layer_load), so reading later layers overlaps with computing earlier ones. Caches the engine reads before any layer hook — the indexer of DeepSeek Sparse Attention, the recurrent state of Mamba-type layers — are restored before the pass.
  4. Save. After a layer computes, the pages it added are copied into the ring and written by the gateway asynchronously (PUT). On decode steps under full CUDA graphs the writes of a step go out as one batched request (deferred write), so the engine keeps its fastest graphs.

The gateway gives the ranks of one engine one shared view: it waits until every rank has sent a request, then runs it once. A restore of a 196,000-token prompt is one request per layer of each rank, not one per page; the ring capacity is checked per request, so it does not limit the context length.

Keys and persistence

A stored object is one chunk of tokens of one layer. Its key is derived from:

  • vLLM’s block hash of the chunk — a hash over the tokens of the chunk and every chunk before it, which already carries the request’s cache_salt, LoRA adapter and multimodal inputs;
  • the layer, and for layers with two caches (DeepSeek Sparse Attention) which of them;
  • optionally a deployment namespace (key_namespace).

Nothing about a key depends on the process that wrote it, so a restarted model server finds everything stored before — as long as PYTHONHASHSEED is fixed: without it Python salts the block hashes per process and nothing is ever found again. Two requests with the same prefix share the stored chunks; two tenants with different cache_salt values never do. Tenant isolation covers the setup.

The chunk (chunk_size, in tokens) is the unit of a stored object and of a hit: a prompt of N tokens can hit at most floor(N / chunk_size) × chunk_size tokens; the tail is recomputed.

WARNING

The layout of the stored data is derived from the configuration, not recorded on the card. After changing chunk_size, key_namespace, the fragment settings of the gateway, the tensor-parallel size or the model, clean the cache: old objects are no longer found, and with a different fragment layout they would not be read back correctly. Settings that require a clean cache.

Fragments

The card accepts objects up to 256 KiB, and over NVMe-oF one command carries at most the device’s maximum transfer size (MDTS, typically 128 KiB). A KV page of a hybrid model is 2–3 MiB. The gateway therefore splits an object larger than fragment_size into several keys and reassembles it on read; the fragments of one object are spread over different databases so they are read in parallel. The fragment size follows from the placement: 131072 for a remote card, 261120 for a local one. See ADP card setup.

The shared-memory ring

The ring is a file in /dev/shm that the gateway creates and every rank maps. Its size (shared_memory_size, 6 GiB by default, 10 GiB is common for 8-GPU servers) sets how many transfers can be in flight:

max_parallel_ios = (shared_memory_size / 2) / io_size

where io_size is one object (the gateway logs it at start). A request that needs more slots than are free runs in batches: the gateway hands out a slot when it is free and takes it back once the page has reached GPU memory. A larger ring helps only when a single request — one layer of a long prompt with a small chunk — needs more slots than the ring has; it costs host memory for as long as the gateway runs. Both containers must see the same /dev/shm.

Deployment topologies

TopologyNotes
One vLLM instance, one gateway, one cardThe basic unit. One gateway serves the ranks of exactly one engine.
Several vLLM instances on one serverOne gateway per instance, each with its own port and ring.
Several deployments on one cardGive each a different key_namespace, so two models of the same geometry or two front ends with different salt policies never read each other’s objects.
Card in a storage serverThe GPU server connects the card’s NVMe-oF namespace over RDMA; the gateway on the GPU server uses it as a local block device.

When the card is unavailable

The external cache never holds generation hostage. If the gateway stops answering, the connector switches the external cache off, logs External KV cache turned off on this rank once, and the engine carries on: an attention-only model serves from GPU memory alone, and a hybrid model (Mamba-type layers) stops cleanly rather than serve a partly restored recurrent state. To bring the external cache back, restart the gateway and vLLM together (the whole pod on Kubernetes).

On MLA models with tensor parallelism every rank needs the same latent page. With nvlink_fanout enabled, each rank of an NVLink group copies only its share of a layer out of host memory and the ranks exchange the rest over NVLink, right before the layer computes. Host-to-GPU traffic of a restore drops from one copy per rank to one per group: in our measurements the copy is 2.14× faster and decode under 32 concurrent agents 14% faster (release notes). It needs NVIDIA GPUs with NVLink peer access and takes a staging buffer and an NCCL communicator outside gpu_memory_utilization. Settings: vLLM KV connector.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs