What it does
A model server keeps the KV cache of every prompt it has seen in GPU memory, and only for as long as that memory lasts. When a long conversation, an agent’s context or a shared system prompt is evicted, or when the server restarts, the next request that starts the same way pays the full prefill again: tens of seconds for a long context on a large model.
The ADP KV cache connector gives vLLM a second, much larger tier for that cache. Every KV page the engine computes is also written to the ADP card; when a prompt comes back, vLLM restores its prefix from the card instead of recomputing it. Pages are keyed by vLLM’s own block hashes, so the cache survives restarts of the model server and is shared by every request with the same prefix. A 196,000-token prompt on GLM-5.3 comes back in 2.1 seconds instead of 47.7 seconds of prefill (release v6.1.1).
What you get:
- Restore instead of recompute — the time to first token of a returning long prompt drops by an order of magnitude, and the GPUs spend that time on new work.
- A cache that outlives the GPU — terabytes of KV cache on NVMe behind the card, kept across restarts, upgrades of the model server and evictions from GPU memory.
- No extra host RAM tier — the only host memory the connector takes is a shared-memory transfer window of a few GiB; layers are restored one at a time while the model computes.
- Plain vLLM — the connector plugs into vLLM’s KV connector API with one command-line option; no fork of the engine.
Architecture
The connector runs inside vLLM, next to every tensor-parallel rank. It hands KV pages to the storage gateway, a separate process on the same server, through a ring in shared memory. The gateway schedules the reads and writes, splits large objects into fragments the card accepts and spreads them over the card’s databases. The ADP card stores them on NVMe SSDs.
The card can sit in either of two places, and nothing above the gateway changes between them:
| Placement | Where the card is | How the gateway reaches it | Gateway backend |
|---|---|---|---|
| Local | a PCIe slot of the GPU server | the ADP software stack on the same host | local-xdp |
| Remote | a storage server | NVMe over Fabrics on an RDMA network (RoCE v2 or InfiniBand) | remote-xdp |
A remote card keeps the GPU servers free of storage and lets the cache sit on hardware sized for it; a local card needs no fabric. Architecture follows a page through both paths.
Compatibility at a glance
| Supported | |
|---|---|
| Inference engine | vLLM 0.26, 0.28 and 0.29 — one connector code base; release images for 0.28.0 and 0.29.0 |
| GPUs | any NVIDIA GPU the vLLM image supports (compute capability 7.5 and newer in the upstream images) |
| CUDA | the CUDA of the vLLM base image: 13.0 for the default vllm/vllm-openai tags, 12.9 for the -cu129 tags |
| Model architectures | full attention, sliding window, MLA, DeepSeek Sparse Attention, Mamba / Mamba-2 / Gated DeltaNet / Kimi Delta Attention hybrids, MoE; multi-token prediction (MTP) |
| Parallelism | tensor parallelism; pipeline parallelism 1 for hybrid models |
| Host | Linux x86_64, Docker or Kubernetes |
The models verified on hardware, with the vLLM version each is recommended on, are listed in Supported models. Full requirements: Requirements.
Integration options
There are three ways to put the ADP card under vLLM. They share the gateway and the card; they differ in how the engine hands its KV cache over. Choosing an integration compares them.
Next steps
- Check the requirements for the GPU server, the card and the network.
- Set up the ADP card, local or remote.
- Install the gateway and vLLM with the connector, and verify a restore.
- Tune with the configuration reference and the notes for your model.
- Serving several tenants from one deployment? Set up tenant isolation.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.