What it is
The plugin makes the ADP card the second tier (L2) of the LMCache multiprocess server. LMCache keeps everything it
does — its interface to vLLM, its index, its policies and its CPU tier (L1, pinned host memory); objects that leave L1,
or are looked up after a restart, are stored on the card behind the storage gateway. LMCache
loads the plugin through its own native_plugin adapter type: LMCache itself is not changed.
Use it when LMCache is already part of your stack. If it is not, the vLLM KV connector restores faster and needs no host-memory tier (comparison).
How it stores data
- Objects and fragments. One LMCache chunk is one object; it is cut into fragments of
fragment_bytes, each one I/O. A small manifest is written only after every fragment of the object has been acknowledged: its presence is what “the object is on the card” means. - Integrity. A load checks the manifest — size, geometry, key — and, by default, a sample checksum of every fragment. Anything that does not match is a miss for that object only, and an unusable manifest is removed so the next store writes the object again.
- Never written twice. Before a store the plugin asks the card whether the object is already there, and skips it if so: a key is a hash of the tokens, so the same chunk is the same entry.
- Failures. A missing key or fragment is a miss. A gateway that stops answering (
wait_timeout_s) turns every answer into a miss until the plugin opens a new session (reopen_after_s) — which is also how a restarted gateway is picked up. A gateway that is not up yet when the server starts is the same case. - Backpressure. Stores beyond
max_pending_store_mbof backlog are refused at once, so L1 objects are not held up behind a slow card; loads take priority over stores.
Requirements
| vLLM | 0.28.0 |
| LMCache | 0.5.4, as bundled with the vllm/vllm-openai:v0.28.0 image, in multiprocess (server) mode |
| Image | vLLM 0.28.0 with the plugin compiled on it, delivered by Awide Labs on request |
| Gateway and card | as for the connector: ADP card setup, the gateway from Installation |
Run it
The LMCache server takes the plugin as its L2 adapter:
lmcache server --host 127.0.0.1 --port 6555 --l1-size-gb 120 --eviction-policy LRU --chunk-size 256 \
--l2-adapter '{"type":"native_plugin","module_path":"xkv_lmcache","class_name":"XkvL2Connector",
"adapter_params":{"port":5557,"fragment_bytes":131072}}'
and vLLM connects to the LMCache server:
--kv-transfer-config '{"kv_connector":"LMCacheMPConnector",
"kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector",
"kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.port":6555}}'
Inside the vLLM container. The image starts the LMCache server from the container’s main process when LMC_ENABLE=1
is set, and waits until its port answers. The server and its L1 then restart with vLLM.
LMC_ENABLE=1 LMC_L1_GB=120 LMC_CHUNK=256 LMC_XKV_PORT=5557 LMC_XKV_FRAG=131072
# optional: LMC_XKV_VERIFY, LMC_XKV_COPY_THREADS, LMC_XKV_STATS_PATH=/tmp/xkv-stats.json,
# LMC_XKV_PARAMS='{"batch_slots":2048}' (merged last); python3 -m xkv_lmcache_autostart --help lists them all
As a separate process, so that L1 survives a restart of vLLM: the same image, started with
python3 -m xkv_lmcache_autostart --exec (without LMC_ENABLE). It needs the GPUs (it copies L1 to the GPU through
CUDA IPC), the host IPC namespace and /dev/shm shared with vLLM and the gateway, and the gateway port:
docker run -d --name lmcache-server --gpus all --network host --ipc host \
-e LMC_HOST=0.0.0.0 -e LMC_PORT=6555 -e LMC_L1_GB=120 -e LMC_CHUNK=256 \
-e LMC_XKV_PORT=5557 -e LMC_XKV_HOST=<gateway address> -e LMC_XKV_FRAG=131072 \
--entrypoint python3 <image> -m xkv_lmcache_autostart --exec
# vLLM: "kv_connector_extra_config": {"lmcache.mp.host": "<server address>", "lmcache.mp.port": 6555}
On Kubernetes that is a second container in the model pod, or a host-network, host-IPC pod on the node.
Settings
The plugin is configured through the adapter_params of the L2 adapter (or the LMC_XKV_* variables of the
autostart); every key, with its default, is in Configuration → LMCache plugin.
Set fragment_bytes to the gateway’s fragment_size. Leave LMCache’s own L2 limit max_capacity_gb at 0: space on
the card is managed by its key eviction.
What to expect
- Hits in L1 come back slightly faster than from the card through the connector, at the cost of the host memory L1 takes. In our comparison on GLM-5.3, a 196,000-token prompt held in a 400 GiB L1 returned in 1.68 s against 1.86 s through the connector (details).
- Hits from the card go card → L1 → GPU: LMCache finishes the whole prefetch into L1 before it copies to the GPU, so a restore after an engine restart is slower than through the connector (2.91 s against 2.26 s in the same comparison). That is inside LMCache, not the plugin.
- Host memory: the data passes through two host buffers, the gateway ring and L1.
Limits: one LMCache server per node and one gateway per server; two servers must not write one card; x86_64 Linux only.
Monitoring
Every stats_interval_s the plugin writes a summary line prefixed [xkv-lmcache] to the LMCache server log;
stats_path writes the full JSON: per-operation tasks, hits, bytes, latency percentiles, queue depths, ring use,
integrity counters, session state and process memory. The store outcomes are told apart by store_skipped (already on
the card), store_deduped, store_dropped (backlog full) and store_rejected_by_card — the last one growing means the
card refuses writes (capacity).
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.