Documentation menu
Integration options

Choosing an integration

Three ways to put the ADP card under vLLM — the vLLM KV connector, the LMCache L2 plugin and the vLLM OffloadingSpec backend — what each one is, and when to use it.

All three integrations store the KV cache on the same card through the same storage gateway; they differ in which part of the stack hands the KV cache over and how it comes back to the GPU.

The three options

Three paths to the same card
what sits between vLLM and the ADP card
vLLM KV connectorGA · RECOMMENDEDvLLM + ADP connectorinside every rankShared-memory ringtransfer windowStorage gatewayfragments, databasesADP cardNVMe SSDs behind itrestores layer by layerLMCache L2 pluginSUPPORTEDLMCache serverL1 in host RAM, under vLLMADP L2 pluginLMCache unchangedStorage gatewayfragments, databasesADP cardNVMe SSDs behind ithost-RAM tier, card as L2OffloadingSpec backendEXPERIMENTALvLLM offloadinghandles every modelADP backendbytes under keysStorage gatewayfragments, databasesADP cardNVMe SSDs behind itany model, any vLLM version

vLLM KV connector — GA, recommended. Our own connector for vLLM’s KV connector API. It restores a prompt layer by layer while the model computes, handles every architecture in the supported models list, and keeps no cache in host memory. This is what the product image ships and what the rest of this documentation describes. Details.

LMCache L2 plugin — Supported. For deployments that already run LMCache. The ADP card becomes the L2 tier of the LMCache multiprocess server, below LMCache’s own CPU tier: LMCache keeps its interface, index and policies, and loads the plugin through its standard native-plugin interface — LMCache itself is not changed. Details.

vLLM OffloadingSpec backend — Experimental. A backend for vLLM’s own KV offloading framework (OffloadingConnector), where vLLM handles the model-specific layout and the backend only stores bytes under keys — so it works with any model and any vLLM version the framework supports, with no per-model work. A prototype for evaluation, not for production yet. Details.

Comparison

vLLM KV connectorLMCache L2 pluginOffloadingSpec backend
StatusGA · recommendedSupportedExperimental
Plugs intovLLM KV connector APILMCache multiprocess server (native_plugin L2 adapter)vLLM offloading framework
vLLM sideFusIOnXConnectorLMCacheMPConnectorOffloadingConnector + XkvOffloadingSpec
Modelsthe verified list; a new architecture needs connector supportwhat LMCache supportsany model vLLM’s offloading framework handles, no per-model work
Restore pathcard → ring → GPU, layer by layer, overlapped with computecard → LMCache CPU tier → GPU; hits in the CPU tier skip the cardcard → GPU staging → GPU, the whole prefix before the step computes
Host memorythe shared-memory ring only (6–10 GiB)LMCache’s CPU tier (sized by you) plus the ringstaging slots plus the ring
vLLM versions0.26, 0.28, 0.290.28 with LMCache 0.5.4any version with the offloading framework (checked against 0.26, 0.28, 0.29)
Delivered asthe product image and the client kitan image of its own, on requeston request, for evaluation

How to choose

  • Start with the vLLM KV connector. It is the fastest path back from the card and the one every model on the supported list is verified with. With the engine restarted, GLM-5.3 gets a 196,000-token prompt back in 2.26 s through the connector and in 2.91 s through LMCache with the ADP card below it (the DRAM tier comparison).
  • Choose the LMCache plugin if LMCache is already part of your serving stack — its CPU tier, its eviction policies, its integrations — and you want the card underneath instead of re-architecting. A prompt that is still in LMCache’s CPU tier comes back slightly faster than from the card (1.68 s against 1.86 s in the same comparison); a prompt that has to come from the card pays for LMCache’s two-step copy, card to host memory, then to the GPU, and the CPU tier costs host RAM.
  • Try the OffloadingSpec backend when model coverage matters most — any model, any vLLM version with the offloading framework — and you are evaluating rather than serving production traffic (current limits).

The gateway, the card setup and the configuration of the gateway are the same for all three.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs