Documentation menu
Security

Tenant isolation

Keeping one tenant's KV cache from another's, on the GPU and on the ADP card: how the salt reaches the storage key, what the API front end has to do, a working front end, the connector's switches, and how to verify it.

Why it matters

Prefix caching shares computed KV blocks between every request that starts the same way, whoever sent it. A hit skips the prefill, and the prefill is most of the time to first token of a long prompt: on GLM-5.3 a prompt of about 3,000 tokens answers in 672 ms from scratch and in 92 ms from the cache. So any client can tell whether someone already sent a given text — guess, time, repeat — and reconstruct another tenant’s prompt block by block. In vLLM this is CVE-2025-46570. An external KV cache widens the window: the cache now outlives GPU memory and restarts.

For a single team or application there is nobody to hide prompts from and nothing to do. For a service shared by several customers, or by users with private data, isolate the cache as described here.

How isolation reaches the card

vLLM accepts a cache_salt field in the request body (chat completions, completions, the Responses API, pooling and the Anthropic-compatible /v1/messages). The salt is mixed into the hash of the first block of the prompt, and every later block hashes its parent, so the salt changes every block hash of the request. Requests with different salts never share a block; requests with the same salt share exactly as before.

The ADP KV cache connector does not hash tokens itself. Every key it writes to the card is built from vLLM’s salted block hashes — in every cache group, including Mamba state, MTP layers and the sparse-attention indexer — so a tenant’s scope holds on the card exactly as it does in GPU memory, with no setting on the connector.

The salt's path into the storage key
set once in the front end, carried to the card as is
Front end authenticates, drops client salts cache_salt = HMAC(K, tenant) vLLM h₁ = H(salt, block 1) h₂ = H(h₁, block 2) … a salted hash chain ADP connector storage key = salted block hash + layer (+ key_namespace) ADP card one key space per salt tenants never share The same path for every cache group: attention, Mamba state, MTP layers, the sparse-attention indexer.

Isolation is only as good as the component that sets the salt. vLLM does not authenticate anyone, accepts whatever salt the client sends and treats an empty salt as none. That component is the API front end in front of vLLM.

1. Choose the scope

A salt splits the cache: a prefix common to two scopes — a long shared system prompt, the same document — is computed and stored once per scope instead of once overall. Inside a scope nothing changes. Put the boundary where trust ends:

DeploymentSaltWhy
One team or one application, one trust domainnoneNobody to hide prompts from; a salt only costs hits.
Shared public system prompt, few-shot examples, a documentation assistantnone or per tenantThe shared prefix is the point of the cache, and there is nothing secret in it.
An API or SaaS serving several customers from one deploymentper tenant, at leastCustomers must not learn each other’s prompts; colleagues still share.
Users of one organization with private data: HR, medical, legal, personal mailper userThe boundary runs between people, not between companies.
One-off prompts with secrets, such as credentials pasted into a chatper session or requestNo reuse at all — use it narrowly, or keep such requests out of the external cache with xkv_store: false.

A client may narrow its scope, never widen it.

2. Set the salt in the front end

Six rules for the API front end:

  1. Set the salt on every request, from the authenticated identity. Drop any cache_salt the client sent, and the fields that reach the engine or its connectors unfiltered — kv_transfer_params, vllm_xargs, and client-supplied multimodal uuid values, which feed the same hash.
  2. Make the salt unguessable. Derive it with HMAC-SHA256 from a secret of at least 32 bytes that only the front end holds. A salt equal to the tenant name or account id can be guessed, and a guessed salt reopens the channel.
  3. Match the scope to the trust boundary (above).
  4. Expose only the endpoints that carry the salt. vLLM must not be reachable around the front end, and endpoints without a salt field stay closed.
  5. Close the internal surfaces — next step.
  6. Keep the salt stable. Derived from the secret and the identity, the same tenant gets the same salt after a restart, and the ADP card keeps serving that scope. Rotating the secret empties the cache for everyone: the right tool after a leak, the wrong one on a schedule.
Where the change goes
one addition in the gateway; vLLM and the KV tier unchanged
Clients tenant A tenant B … API gateway auth: key → tenant + salt middleware drop client salt cache_salt = HMAC(K, tenant) vLLM salt enters every block hash KV cache tier connector keys = salted block hashes, as is no direct access to vLLM INTERNAL NETWORK ONLY /metrics · /load · KV events · storage ports of the KV tier THE ONLY NEW CODE IS THE PURPLE BOX · VLLM AND THE KV TIER STAY AS THEY ARE

A front end in Python with FastAPI, for chat and completion requests, streaming included. The same few lines go into any proxy you already run.

import base64, hashlib, hmac, os

import httpx
from fastapi import FastAPI, HTTPException, Request
from fastapi.responses import JSONResponse, StreamingResponse
from starlette.background import BackgroundTask

SECRET = bytes.fromhex(os.environ["CACHE_SALT_SECRET"])   # 32+ random bytes, only the front end has them
DEPLOYMENT = "prod"                                        # part of the salt: two deployments never share
STRIP = ("cache_salt", "kv_transfer_params", "vllm_xargs") # never taken from the client

app = FastAPI()
vllm = httpx.AsyncClient(base_url="http://vllm:8000", timeout=None)


def cache_salt(scope: str, scope_id: str) -> str:
    """The same scope always gets the same salt; nobody without SECRET can compute it."""
    message = f"v1|{DEPLOYMENT}|{scope}|{scope_id}".encode()
    digest = hmac.new(SECRET, message, hashlib.sha256).digest()
    return base64.urlsafe_b64encode(digest).rstrip(b"=").decode()


def authenticate(request: Request) -> str:
    """Your existing authentication: API key -> tenant id."""
    tenant = lookup_tenant(request.headers.get("authorization", ""))
    if tenant is None:
        raise HTTPException(status_code=401)
    return tenant


def salted(body: dict, tenant: str) -> dict:
    for field in STRIP:
        body.pop(field, None)
    for message in body.get("messages", []):               # client-supplied multimodal ids feed the hash too
        if isinstance(message.get("content"), list):
            for part in message["content"]:
                if isinstance(part, dict):
                    part.pop("uuid", None)
    body["cache_salt"] = cache_salt("tenant", tenant)      # or ("user", user_id) for a per-user scope
    return body


async def forward(request: Request, path: str):
    body = salted(await request.json(), authenticate(request))
    if body.get("stream"):
        upstream = await vllm.send(vllm.build_request("POST", path, json=body), stream=True)
        return StreamingResponse(upstream.aiter_raw(), status_code=upstream.status_code,
                                 media_type=upstream.headers.get("content-type"),
                                 background=BackgroundTask(upstream.aclose))
    upstream = await vllm.post(path, json=body)
    return JSONResponse(upstream.json(), status_code=upstream.status_code)


@app.post("/v1/chat/completions")
async def chat_completions(request: Request):
    return await forward(request, "/v1/chat/completions")


@app.post("/v1/completions")
async def completions(request: Request):
    return await forward(request, "/v1/completions")

Generate the secret once (openssl rand -hex 32), keep it in your secret store, and give it to the front end only.

3. Close the internal surfaces

SurfaceWhyWhat to do
vLLM’s API portRequests that bypass the front end carry no salt, or one the client chose.Reachable from the front end only.
/metrics, /loadNot covered by the API key; they count hits globally.Internal network only.
KV events streamCarries tokens and salts in the clear.Internal network only, or off.
The gateway’s control portAnswers lookups and deletes without authentication.--bind-address=127.0.0.1 when vLLM shares its network namespace; never on a tenant network.
The NVMe-oF fabric of a remote cardStorage traffic.A storage network, not reachable from tenants.

4. Turn on the connector’s switches

Isolation works without any connector setting. Two switches make it stricter, and one request field keeps a request out of the card entirely. All are off by default, and with them off the storage keys are byte for byte those of a deployment without isolation — so an upgrade needs no wipe.

SettingUse it when
require_cache_salt: trueYour front end salts every request. A request it missed — no salt or an empty one — then neither reads nor stores the external cache (one warning on the first), instead of sharing its prompt with every other unsalted request.
key_namespace: "<deployment>"Several deployments share a card and could present the same block hashes: two models of the same geometry, two front ends with different salt policies. Up to 64 bytes. Setting or changing it starts a new cache.
"kv_transfer_params": {"xkv_store": false} in a requestRequest-scoped traffic whose salt nobody will match again: no lookup, no restore, no write; nothing derived from the prompt reaches the card. Set by the front end only.
--kv-transfer-config '{"kv_connector":"FusIOnXConnector",
  "kv_connector_module_path":"xkv_vllm_connector.fusionx_connector","kv_role":"kv_both",
  "kv_connector_extra_config":{"xkv_gateway_port":5557,"chunk_size":128,
    "require_cache_salt":true,"key_namespace":"prod-glm53"}}'

5. Verify it

Send the same long prompt through the front end as two tenants, before and after a restart of vLLM. Start vLLM with --enable-prompt-tokens-details so responses report cached_tokens.

FRONT=http://front-end.internal:8080
PROMPT=$(python3 -c "print(' '.join(f'Line {i}: a private document of tenant A.' for i in range(3000)))")
ask() {   # ask <api key>: cached prompt tokens of one request
  curl -s "$FRONT/v1/completions" -H "Authorization: Bearer $1" -H 'Content-Type: application/json' \
    -d "$(jq -n --arg p "$PROMPT" '{model: "glm-5.3", prompt: $p, max_tokens: 4}')" \
    | jq '.usage.prompt_tokens_details.cached_tokens'
}

ask "$KEY_TENANT_A"     # 0: the first time anyone sends it
ask "$KEY_TENANT_A"     # close to the prompt length: same tenant, same salt
ask "$KEY_TENANT_B"     # 0: another tenant, another key space
# restart vLLM (not the gateway), wait for /health, then:
ask "$KEY_TENANT_A"     # close to the prompt length: restored from the ADP card
ask "$KEY_TENANT_B"     # still 0

A non-zero answer for tenant B means a request reached vLLM without the front end’s salt, or with a salt shared between the tenants.

Per-tenant salt: every probe looks the same
measured on GLM-5.3 FP8 · salt set by the front end
Tenant A · victim sends prompt P Tenant A · colleague sends P again Tenant B · attacker sends guesses G Front end drops client salts, sets cache_salt from the tenant's identity s_A = HMAC(K, A) s_B = HMAC(K, B) vLLM prefix cache h₁ = H(salt, block tokens) KEY SPACE OF A h₁(s_A,P) h₂ h₃ KEY SPACE OF B h₁(s_B,G) h₂ h₃ hit miss WHAT THE ATTACKER MEASURES guess G ≠ P · miss full prefill 679 ms guess G = P · miss too other salt, other hashes 676 ms no signal TTFT, MEDIAN OF 30 PROBES · GLM-5.3 FP8, ~3,000-TOKEN PROMPT, PER-TENANT SALT

That is what we measured on GLM-5.3 with the ADP KV cache connector, with the GPU prefix cache reset between probes so that any hit had to come from the card: without a salt, the attacker’s probe of the victim’s prompt answered in 92 ms against 672 ms for a prompt nobody sent; with a per-tenant salt, 676 ms against 679 ms — no signal. A colleague in the victim’s tenant kept every cached token, and after a restart of vLLM the card still served that tenant and no other. The full study: Isolating the KV cache between tenants.

What a salt does not cover

  • Inside a scope everything is shared by design, including the cached-token counters. Make the scope no wider than the trust boundary.
  • Load and eviction. Tenants still share GPU memory and the card’s capacity. A tenant can notice that the system is busy or push others’ blocks out of the cache, but cannot learn what their prompts contain.
  • Other channels. Speculative decoding leaks through the number of tokens accepted per step, streaming through packet sizes and timing, a cache-aware scheduler through queue order. These need their own measures.
  • The operator. Whoever runs the hosts, holds the front end’s secret or reads the KV events stream is outside this model.

Planning a deployment?

Talk to our engineers about your models, context lengths, concurrency and where the ADP card should sit.

Talk to Awide Labs