All three integrations store the KV cache on the same card through the same storage gateway; they differ in which part of the stack hands the KV cache over and how it comes back to the GPU.
The three options
vLLM KV connector — GA, recommended. Our own connector for vLLM’s KV connector API. It restores a prompt layer by layer while the model computes, handles every architecture in the supported models list, and keeps no cache in host memory. This is what the product image ships and what the rest of this documentation describes. Details.
LMCache L2 plugin — Supported. For deployments that already run LMCache. The ADP card becomes the L2 tier of the LMCache multiprocess server, below LMCache’s own CPU tier: LMCache keeps its interface, index and policies, and loads the plugin through its standard native-plugin interface — LMCache itself is not changed. Details.
vLLM OffloadingSpec backend — Experimental. A backend for vLLM’s own KV offloading framework
(OffloadingConnector), where vLLM handles the model-specific layout and the backend only stores bytes under keys —
so it works with any model and any vLLM version the framework supports, with no per-model work. A prototype for
evaluation, not for production yet. Details.
Comparison
| vLLM KV connector | LMCache L2 plugin | OffloadingSpec backend | |
|---|---|---|---|
| Status | GA · recommended | Supported | Experimental |
| Plugs into | vLLM KV connector API | LMCache multiprocess server (native_plugin L2 adapter) | vLLM offloading framework |
| vLLM side | FusIOnXConnector | LMCacheMPConnector | OffloadingConnector + XkvOffloadingSpec |
| Models | the verified list; a new architecture needs connector support | what LMCache supports | any model vLLM’s offloading framework handles, no per-model work |
| Restore path | card → ring → GPU, layer by layer, overlapped with compute | card → LMCache CPU tier → GPU; hits in the CPU tier skip the card | card → GPU staging → GPU, the whole prefix before the step computes |
| Host memory | the shared-memory ring only (6–10 GiB) | LMCache’s CPU tier (sized by you) plus the ring | staging slots plus the ring |
| vLLM versions | 0.26, 0.28, 0.29 | 0.28 with LMCache 0.5.4 | any version with the offloading framework (checked against 0.26, 0.28, 0.29) |
| Delivered as | the product image and the client kit | an image of its own, on request | on request, for evaluation |
How to choose
- Start with the vLLM KV connector. It is the fastest path back from the card and the one every model on the supported list is verified with. With the engine restarted, GLM-5.3 gets a 196,000-token prompt back in 2.26 s through the connector and in 2.91 s through LMCache with the ADP card below it (the DRAM tier comparison).
- Choose the LMCache plugin if LMCache is already part of your serving stack — its CPU tier, its eviction policies, its integrations — and you want the card underneath instead of re-architecting. A prompt that is still in LMCache’s CPU tier comes back slightly faster than from the card (1.68 s against 1.86 s in the same comparison); a prompt that has to come from the card pays for LMCache’s two-step copy, card to host memory, then to the GPU, and the CPU tier costs host RAM.
- Try the OffloadingSpec backend when model coverage matters most — any model, any vLLM version with the offloading framework — and you are evaluating rather than serving production traffic (current limits).
The gateway, the card setup and the configuration of the gateway are the same for all three.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.