Where settings live
| Component | Set in | Reference |
|---|---|---|
| Storage gateway | command-line flags, or a YAML file given with --config | Storage gateway |
| vLLM KV connector | kv_connector_extra_config of --kv-transfer-config; per request kv_transfer_params | Connector |
| vLLM process | environment variables | Environment variables |
| vLLM | its command line | vLLM options |
| ADP storage service | /etc/pliops/<N>/pliostore.ini on the card’s host | ADP storage service |
| LMCache plugin | adapter_params of the L2 adapter | LMCache plugin |
| OffloadingSpec backend | xkv.* keys in kv_connector_extra_config | OffloadingSpec backend |
Complete deployments — compose files and a Kubernetes pod — are in Installation and Kubernetes.
Storage gateway
The gateway takes its settings from the command line, from a YAML file (--config <file>), or both; a flag given on
the command line wins over the file. The flag is the key with dashes: shared_memory_size → --shared-memory-size.
Sizes take B, KiB, MiB, GiB, TiB (1024-based) or KB, MB, GB, TB (1000-based); a bare number is
bytes.
# gateway.yaml — a remote card on an 8-GPU server
storage_backend: remote-xdp
storage_devices:
- /dev/nvme1n1
fragment_size: 131072
port: 5557
shared_memory_size: 10GiB
thread_pool_size: 64
num_dbs: 128
start_clean: 0
log_severity: info
| Key | Default | What it does |
|---|---|---|
storage_backend | remote-xdp | Where objects go: remote-xdp (an ADP card over NVMe-oF), local-xdp (an ADP card in this server), file-system (a directory; for trying the software without a card), loopback (stores nothing; for tests). |
storage_devices | — | Required for remote-xdp: the connected NVMe-oF namespace(s) of the card, for example /dev/nvme1n1. The gateway refuses to start without it. |
fragment_size | 0 (off); 261120 on local-xdp | Objects larger than this are stored as several fragments. Local card: 261120. Remote card: a multiple of 4 KiB no larger than the device’s transfer limit — 131072. At least 4096 when set. ⚠️ Changing it requires a clean cache. |
fragment_spread_across_dbs | true | Put the fragments of one object into different databases, so they are read in parallel. ⚠️ Changing it requires a clean cache. |
port | 5557 | The control port (ZeroMQ). Also names the ring file in /dev/shm, so gateways on one host need different ports. |
bind_address | * | The address the port listens on: an IPv4 address, an interface name or *. The port answers lookups and deletes without authentication — bind it to 127.0.0.1 when vLLM shares the gateway’s network namespace (one pod, or host networking), and keep it off tenant networks otherwise. |
shared_memory_size | 6GiB (at least 512MiB) | The ring in /dev/shm, split between the tensor-parallel ranks of the engine. A transfer window, not a cache: its capacity is checked per request, so it does not limit the context length. 10 GiB is a good value for 8-GPU servers. |
thread_pool_size | 4 | Worker threads; each runs one device command at a time, so this is also the most commands in flight. The main lever of restore speed: 64 on 8-GPU servers. |
core_list | all cores | CPU cores to pin the workers to, for example "0-7,16-23". |
num_dbs | 128 | Number of databases on the card the objects are spread over. |
start_clean | 0 | 1 deletes the gateway’s databases at start, i.e. wipes the cache. Use once, deliberately; keep 0 while key eviction is enabled on the card (wiping the card). |
scheduler_max_parallel_reqs | 2 | Read requests running at once. Keep 2, the value validated on hardware (the gateway warns above it); long restores get faster with thread_pool_size, not with this. |
scheduler_overlap_trigger_ratio | 0.8 (1.0 on local-xdp) | Share of the previous read that must be done before the next one may start. |
sequence_timeout_sec | 30 (at most 33) | How long a request sent by some ranks of an engine may wait for the others before it fails. |
log_dir | empty | A directory for log files; empty logs to standard output only. |
log_severity | info | perf, trace, debug, info, notice, warning, error or fatal. |
Connector
The keys of kv_connector_extra_config in --kv-transfer-config. A boolean key also takes 1/0, "true"/"false",
"yes"/"no", "on"/"off"; anything else stops the connector at start.
| Key | Default | What it does |
|---|---|---|
xkv_gateway_port | 5557 | The gateway’s control port. |
chunk_size | vLLM’s scheduler block size | Tokens per stored object; must be a multiple of the block size of every KV cache group (for sliding-window models the window must be at least twice chunk_size and a multiple of it). ⚠️ Changing it starts a new cache. |
slot_alignment_bytes | 0 | Round each transfer up to this many bytes. 4096 on a remote card for models whose page is not a multiple of 4 KiB (GLM-5.x). ⚠️ Changing it requires a clean cache. |
enable_deferred_write | true | Batch the writes of decode steps that run under full CUDA graphs into one request per step. false makes the connector ask vLLM for piecewise graphs (slower decode). |
layer_callbacks | auto | Whether the engine’s layer hooks are used for this model; auto decides from the model and checks it against the registered layers. |
aggregated_get_submission, aggregated_put_completion | true | Submit reads and complete writes in batches. |
synchronous_layer_get | false | Make every layer wait for its own read before computing even where the read path does not need it. A debugging guard. |
verify_all_layers | false | Confirm a cached prefix on every exchanged layer instead of one layer per group. |
clean_cache | false | Wipe the cache when the connector starts. For tests. |
log_level | INFO | Ranks other than 0 log from NOTICE up unless a level is set. |
num_copy_thread_block_get, num_copy_threads_per_block_get, num_copy_thread_block_put, num_copy_threads_per_block_put | SDK defaults | Copy kernel launch parameters. |
nvlink_fanout | false | Fan restores of MLA models out over NVLink (details). |
nvlink_fanout_groups | auto | NVLink groups of ranks, auto from NVML or explicit. |
nvlink_fanout_staging_mb | 256 | Staging buffer per rank, MiB. |
nvlink_fanout_min_layer_bytes | 4194304 | Smaller layers are copied whole by every rank. |
key_namespace | empty | Mix a deployment name (up to 64 bytes) into every key. ⚠️ Setting, changing or clearing it starts a new cache. |
require_cache_salt | false | Requests without a cache_salt neither read nor store the external cache. |
cache_rules | — | Experimental override of who writes a cache, by layer class. Not for production. |
Per request, in the request body: "kv_transfer_params": {"xkv_store": false} keeps the request out of the external
cache — no lookup, no restore, no write.
Environment variables
| Variable | Where | What it does |
|---|---|---|
PYTHONHASHSEED=0 | vLLM | Required. A fixed hash seed, so block hashes — and the keys on the card — are the same after a restart. |
XKV_GATEWAY_HOSTNAME | vLLM | Host of the gateway’s control port; localhost by default. In Docker Compose, the gateway’s service name. |
VLLM_USE_V2_MODEL_RUNNER=0 | vLLM 0.29 | Required on 0.29: the connector works with the V1 model runner. |
NCCL_LAUNCH_ORDER_IMPLICIT=1 | vLLM | Recommended with NVLink fan-out (NCCL 2.26 or newer); harmless otherwise. |
XKV_EXIST_REOPEN_COOLDOWN_S | vLLM | Seconds between attempts to reopen the lookup connection after a timeout; 30 by default. |
XKV_STEP_METRICS_EVERY | vLLM | Report per-step timing of the connector every this many steps; 500 by default. |
XKV_SYNCHRONOUS_LAYER_GET | vLLM | Overrides synchronous_layer_get. |
XKV_NVTX | vLLM | NVTX ranges around the connector’s work, for profiling. |
vLLM options the connector depends on
| Option | When | Why |
|---|---|---|
--enable-prefix-caching | always | The connector serves prefix-cache hits; hybrid models refuse to start without it. |
--enable-prompt-tokens-details | recommended | Reports cached_tokens in API responses. |
--mamba-cache-mode align | Mamba / GDN / KDA hybrids | The only Mamba cache mode the connector accepts; vLLM 0.29 picks it by itself with prefix caching. |
--watermark=0.02 | Mamba / GDN / KDA hybrids | Keeps the few free blocks a hybrid needs when a request’s prefix is restored, on a nearly full KV cache (details). |
--kv-cache-dtype fp8 | DeepSeek Sparse Attention models, MiniMax M3 | One page size per group. |
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' | DeepSeek Sparse Attention models on vLLM 0.29 | Under piecewise graphs the engine does not write their KV. |
--max-num-batched-tokens=8192 | GLM-5.x | Avoids an out-of-memory error in the sparse MLA kernel on prompts from about 128k tokens. |
--no-async-scheduling | full + sliding-window models (Gemma 4, GLM-4.7-Flash) with deferred write | Deferred write of a sliding-window group needs synchronous scheduling and pipeline parallelism 1; the alternative is enable_deferred_write: false. The connector says so at start. |
--gpu-memory-utilization | NVLink fan-out | About 0.005 lower on a 141 GB GPU for the fan-out buffers. |
ADP storage service
/etc/pliops/<N>/pliostore.ini on the host of card N, written by the installer. Apply a change with the gateway and
vLLM stopped: systemctl restart pliostore@<N>, and on a storage server the target after it
(ADP card setup).
| Section | Key | Value | Why |
|---|---|---|---|
[db_settings] | key_eviction | true | Required. Lets the card reclaim space on its own. Without it a full card refuses every write: vLLM keeps serving, but nothing new reaches the cache and the hit rate decays to zero. It is off unless set. |
[db_settings] | product_type | 0 | Key-value mode, which the gateway uses (1 is block mode). |
[qos] | aio_read_io_size, aio_write_io_size | as delivered (131072) | The I/O size of the service; matches the 128 KiB transfer of the NVMe-oF target. |
[media], [raid], [data], [compression] | — | as delivered | Written by the installer for this card and its SSDs. |
[debug] | — | as delivered | Internal; do not change. |
LMCache plugin
adapter_params of the native_plugin L2 adapter in lmcache server --l2-adapter '{…}'
(LMCache L2 plugin). All keys are optional; an unknown key is refused with the list of
accepted ones.
| Key | Default | What it does |
|---|---|---|
port, gateway_host | 5557, localhost | The gateway. |
fragment_bytes | 131072 | I/O size and fragment width: at least 4096 and at most the device’s transfer limit. Set it equal to the gateway’s fragment_size when that is set (261120 on a local card). Changing it makes everything stored a miss. |
verify | sample | off, sample (first and last 4 KiB of every fragment) or full. |
batch_slots, get_inflight, put_inflight, exist_batch_keys | 1024, 4, 2, 8192 | Request size in ring slots and pipelining. A request never takes more than half the ring. |
copy_threads, copy_chunk_frags | 8, 32 | Copy and checksum workers, and their job size in fragments. |
max_pending_store_mb | 16384 | Store backlog above which new stores are refused; 0 = no limit. |
wait_timeout_s, reopen_after_s | 30, 30 | How long the gateway may stay silent before the session is given up; the delay before a new one (0 = never reopen). |
namespace | empty | Mixed into every key: separate deployments on one card. Changing it makes everything stored a miss. |
start_clean | false | Wipe the card when the plugin opens — for tests only. |
prefault | true | Touch the ring pages at open. |
stats_interval_s, stats_path, log_task_min_objects, log_level | 60, empty, 64, info | Observability: a summary line in the server log every stats_interval_s, JSON statistics to a file. |
Leave LMCache’s own L2 limit max_capacity_gb at 0: space on the card is managed by the card’s
key eviction.
OffloadingSpec backend
xkv.<key> entries (or a nested xkv object) in kv_connector_extra_config; the environment variable
XKV_OFFLOAD_PARAMS, a JSON object, is merged last (OffloadingSpec backend).
Key (xkv. prefix) | Default | What it does |
|---|---|---|
port | 5557 | The gateway. |
fragment_bytes | 131072 | I/O and fragment size; 261120 for a local card, at most the transfer limit of a remote one. |
staging_bytes, staging_min_slots | 2 GiB, 4 | The pool of staging slots. |
pipeline_depth | 2 | Waves of transfers in flight at once. |
max_objects_per_wave | 0 | 0 = the pool divided by pipeline_depth. |
device_staging | auto | auto: slots in GPU memory, one host copy; off: two copies through pinned host memory. |
tp_layout | auto | auto / replicated: one object per tensor-parallel group for MLA pages; per_rank: one per rank (with port_per_rank: true). |
port_per_rank | false | A gateway per rank, for per_rank. |
sync_lookup, sync_lookup_timeout_ms | false, 200 | Probe a new request’s keys inside the scheduler step instead of deferring the request by a step. |
trust_own_stores | false | Whether the scheduler counts its own stores as present without asking the card. |
probe_batch_keys, absent_ttl_s, max_index_entries | 4096, 60, 1000000 | The index of what the card holds. |
op_timeout_s, acquire_timeout_s | 120, 30 | Timeouts of an operation and of acquiring a slot. |
draft_groups | auto | Which KV groups of a hybrid model with MTP count as draft groups: auto keeps the Mamba groups out of them (needed for hits after a restart); vllm uses vLLM’s own marking. |
flush_order | auto | auto completes pending stores before model runner V2 reuses their blocks; vllm keeps vLLM’s order. |
timing | false | A log line per wave and per job. |
Settings that require a clean cache
The layout of what is stored is derived from the configuration and not recorded on the card. After any of these changes, wipe the card (how):
| Change | What happens to the old objects |
|---|---|
fragment_size, fragment_spread_across_dbs (gateway) | Not read back correctly (they are not reported as misses) — a wipe is mandatory. |
slot_alignment_bytes (connector) | The object size changes; old objects are unusable. |
chunk_size, key_namespace (connector) | Never found again; they only take space until evicted. |
| Tensor-parallel size, the model or its revision | Never found again. |
| A release whose notes say the stored layout changed | As the release notes say; wipe. |
A change of PYTHONHASHSEED or of the vLLM version can change the block hashes as well; the cache then refills, and
nothing wrong is ever read.
Planning a deployment?
Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.