KV Cache Connector#
The KV Cache Connector is a flexible interface in TensorRT-LLM that enables remote or external access to the Key-Value (KV) cache. It allows developers to implement custom logic for loading, saving, and managing KV cache blocks, extending the capabilities of the standard KV cache manager.
This document explains the KV Cache Connector architecture, common use cases, and provides a detailed walkthrough of the included example.
Use Cases#
The KV Cache Connector is designed to support a variety of advanced serving scenarios:
KV Cache Offloading: Move KV cache blocks from GPU memory to cheaper/larger storage (CPU RAM, NVMe SSD, or network storage) when they are not immediately needed, and reload them when required.
Custom Disaggregated Serving: Separate the prefill (context processing) and decode (token generation) phases onto different instances or machines. The connector can be used to transmit the KV cache generated during prefill to the decode instances.
KV Cache Sharing / P2P Transfer: Share KV cache states between different model instances or across peer-to-peer connections.
Architecture#
The connector architecture is split into two main components:
Scheduler (Leader): Responsible for orchestration. It decides what needs to be loaded or saved and builds metadata instructions. It runs only on the leader rank (rank 0).
Worker: Responsible for execution. It receives metadata from the scheduler and performs the actual data transfers (loading/saving) on the KV cache tensors. It runs on all ranks.
API Reference#
To implement a custom connector, you must subclass KvCacheConnectorScheduler and KvCacheConnectorWorker.
1. Scheduler (Leader) Interface (KvCacheConnectorScheduler)#
These methods run on the leader process and drive the connector’s behavior.
build_connector_meta(self, scheduler_output: SchedulerOutput) -> objectDescription: The core orchestration method. Called during the scheduling phase. It examines the current requests and decides which blocks need to be loaded from or saved to the external store.
Arguments:
scheduler_outputcontains information about new requests, blocks allocated, current request states, and the cumulativeRequestData.block_hasheschain.block_hashesis read directly from each KV cache block’s stored hash, which the KV cache manager commits as soon as a block becomes full – the value matches the hash that KV cache events will subsequently emit for the same block. The chain only covers beam 0; the executor rejectskv_connector_configat startup whenmax_beam_width > 1, so connectors may assume beam-width-1 inputs.Returns: An arbitrary metadata object (picklable) that describes the tasks for the workers. This object is broadcasted to all workers.
get_num_new_matched_tokens(self, request: LlmRequest, num_computed_tokens: int) -> tuple[int, bool]Description: Called when a new request arrives. It checks to see if any KV cache can be loaded from an external KV store.
Returns: A tuple
(num_tokens, is_async).num_tokensis the number of tokens found in the external cache.is_asyncindicates if the loading will happen asynchronously (background) or requires blocking.
request_finished(self, request: LlmRequest, cache_block_ids: list[int]) -> boolDescription: Called when a request completes generation.
Returns: A boolean indicating if an asynchronous save operation is underway. If
True, the system waits for the operation to complete before releasing the KV cache blocks.Note: under sliding-window attention
cache_block_idscovers the live window, not the whole prompt. See What a connector can persist under a sliding window.
update_state_after_alloc(self, request: LlmRequest, block_ids: list[int])Description: a callback to update internal state after KV cache blocks have been allocated for the prefill.
Note: with chunked prefill,
block_idscovers only the blocks allocated for the first chunk. The remaining blocks arrive as append-deltas inRequestData.new_block_idson later calls tobuild_connector_meta, on the entries underscheduler_output.cached_requests. A connector that treats this callback as its only source of block ids will under-plan. Read both lists:def build_connector_meta(self, scheduler_output): for req in scheduler_output.new_requests: # the first chunk self._plan(req.request_id, req.new_block_ids) for req in scheduler_output.cached_requests: # every later chunk self._plan(req.request_id, req.new_block_ids)
Both example connectors walk
new_requestsonly, so neither one demonstrates this.
update_state_after_alloc_by_layer_group(self, request: LlmRequest, block_ids_by_layer_group: list[list[int]])request_finished_by_layer_group(self, request: LlmRequest, cache_block_ids_by_layer_group: list[list[int]]) -> boolDescription: the per-layer-group forms of the two callbacks above, indexed by layer group id. Entry
[g][i]is the page slot of block ordinaliin layer groupg.When they are called: whenever the KV cache reports page indices per layer group. A page index is scoped to its group — one group per attention window size — so a cache with more than one group can be described no other way, and the flat
block_ids/cache_block_idsare empty there. With a single layer group the flat lists carry that group’s indices as well, so a connector that implements only the flat forms keeps working on those models.Which form to implement: implement exactly one complete set.
Set
Models it covers
per-layer-group
every model, VSWA included
flat
non-VSWA, non-hybrid only
A set is complete when both of its methods are defined. Mixing the two — one method from each — is rejected during executor bring-up, naming the method that is missing. An existing flat connector needs no change on the models it already covers: with a single layer group the base per-layer-group implementation folds back to the flat call. Hybrid / linear-attention models are not enabled for the connector yet; the per-layer-group set is the shape their cache will need. See Running under VSWA.
Running under VSWA#
Under variable sliding-window attention the KV cache allocates one pool per attention window size, and a page index only means something inside its own layer group. A single tensor and a single flat block list cannot describe that, so three methods have to be implemented together:
Method |
Replaces |
|---|---|
|
|
|
|
|
|
Implementing the per-layer-group form of a pair is enough — the flat method it replaces does not also have to be defined.
All three are checked during executor bring-up, before any request is admitted. register_kv_cache_layout refuses there, naming the group and region counts it could not describe, and the two scheduler methods are checked alongside it. Nothing is deferred to the first request, so a partial implementation costs a start-up failure rather than one after the model is loaded.
examples/llm-api/llm_kv_cache_connector_vswa.py is a worked connector for this case. A VSWA connector has to do five things:
Address pages per group.
layout.groups[g].regions[r]gives the byte ranges;region.slot_tensor(i)is page slotiof that group.layout.group_of_layer(layer_id)maps a model layer back to its group, which is what the per-layerwait_for_layer_load/save_kv_layerhooks need.Read the per-group block lists.
RequestData.new_block_ids_by_layer_group[g]carries the page slots; the flatnew_block_idsis empty.Carry the layer group in the cache key, and in every transfer target. This one is a correctness requirement, not a convenience. The same token range exists in every layer group holding different KV, so a key derived from the token sequence alone collides across groups and one group’s bytes will overwrite another’s — then be loaded back into the wrong group. Mix
layer_group_id(or the window size, or the layer set) into the identifier, and carry(layer_group_id, page_slot)rather thanpage_slotalone as the transfer target.Filter out-of-window blocks through
valid_page_slots, and size the store for the window rather than the prompt. See below.Serve a block only when every group holds it. A full-attention group keeps the whole prompt while a sliding group keeps only its window, so the prefix that can be served back is bounded by the smallest window. Stop the lookup at the first block ordinal any group misses.
KvCacheLayout reference#
from tensorrt_llm._torch.pyexecutor.connectors.kv_cache_layout import KvCacheLayout
The type is passed to register_kv_cache_layout; importing it is only needed for a type annotation.
A layout describes the byte ranges that repeat once per page slot. It describes ranges rather than
implying them, which is what lets one type cover MLA (a pool simply has no value buffer), block
scales, sliding-window attention and hybrid models without any of them being a special case.
Attribute |
Meaning |
|---|---|
|
Tokens covered by one page. |
|
Element type of the KV data, for a typed view over a region. |
|
The layer groups, as |
|
One group by id. Raises |
|
The group owning a model layer — what the per-layer hooks route on. |
|
The |
Each KvCacheLayerGroupLayout:
Attribute |
Meaning |
|---|---|
|
The index page slots are scoped to. Dense, starting at 0. |
|
Attention window for the group, or |
|
Global model layer indices in the group — the same index space |
|
The |
|
Total bytes the group occupies for one page slot. |
Each KvCacheRegion is a contiguous byte range that repeats once per page slot:
Attribute |
Meaning |
|---|---|
|
Device address of page slot 0. |
|
Bytes the region covers within one slot. |
|
Distance between consecutive slots. |
|
Number of page slots. |
|
The |
|
|
|
A strided |
|
The bytes of one page slot, raising |
size is not necessarily stride: a region covers one run of adjacent buffers within a slot, and a
slot may hold several runs. For a model with uniform layer shapes the buffers coalesce into a single
region spanning the whole slot, which is the whole-page transfer. A group with more than one region
must be addressed region by region.
The addresses are device addresses, and they stay valid because every cache tier below GPU is rejected at bring-up while a connector is attached. See KV cache tiers.
KV cache tiers are GPU-only under a connector#
A connector registers device addresses and holds them across iterations. Evicting a page to another tier reassigns its GPU slot underneath the connector, so with a connector attached:
Setting
KvCacheConfig.host_cache_sizeorKvCacheConfig.disk_cache_sizeabove zero fails at bring-up, with a message naming both settings.Leaving
host_cache_sizeunset drops the host tier rather than failing, with a log line. That tier is provisioned automatically only to give theMAX_UTILIZATIONscheduler somewhere to spill to via suspend/resume, which a connector run does not use.enable_kv_pool_rebalanceis ignored — startup and inference continue, and the rebalance simply never runs. Rebalance suspends every active request and runs a defragmenting migration that reassigns the same page slots a tier eviction would.
The practical consequence is that a KV-exhausted connector deployment has no secondary tier to
fall back on. The remedies are kv_cache_config.max_tokens,
kv_cache_config.free_gpu_memory_fraction, or lowering max_num_tokens to hand memory back to the
KV pool; the scheduler’s exhaustion error says so directly when a connector is attached.
Block reuse alongside the connector#
Specify
KvCacheConfig.enable_block_reuse=Truealongside a connector. Without it the connector’s prefix is never honoured: either the combination is rejected at start-up, or the lookup, the reads and the device copies are performed and discarded, at no correctness cost but at full latency cost.The start-up check reads the value the cache resolved, not the one you passed: some quantization algorithms, some SM versions and hybrid linear models turn block reuse off on their own, so this error can appear without the flag being set anywhere in your configuration.
A page slot must not be reassigned underneath the connector#
The connector holds page indices across iterations, and RequestData reports only the pages appended
since the last call. Anything that hands a slot the connector already knows about to a different
request therefore goes unreported and corrupts the next transfer against it. Three configurations do
that, and each is rejected at bring-up.
Configuration |
Mechanism |
|---|---|
Speculative decoding |
Rejected draft tokens shrink a request’s page list, and the freed slot goes to whichever request allocates next. The connector is never told the tail block moved. |
A capacity scheduler policy other than |
A destroyed-and-replayed request comes back on different pages, and the connector’s per-request block delta is then measured against pages that were freed with it. |
A host or disk cache tier |
Tier eviction reassigns the GPU slot. See KV cache tiers are GPU-only under a connector. |
The exact set the runtime refuses depends on your cache configuration; the bring-up error is authoritative. Speculative decoding is unsupported with a connector wherever it is not refused — the same page-list shrink happens there.
RequestData fields that may not be populated#
Two RequestData fields can be reported empty depending on the cache configuration. A connector must tolerate both.
Field |
When empty |
Consequence |
|---|---|---|
|
|
No block-hash accessor exists on this path. Nothing in the runtime reads the field, and neither example connector uses it — both hash the token sequence themselves. A connector that keys its external store on |
|
|
|
Both are gaps to be closed rather than intended behaviour.
Blocks with no page#
A block that has no page in a layer group is reported as -1 (BAD_PAGE_INDEX) in place, not dropped from the list. This keeps each entry aligned with its block ordinal, so entry i always describes tokens [i * tokens_per_block, (i+1) * tokens_per_block) and an append-delta over successive calls stays valid.
That alignment is also why the list is not safe to index with directly: -1 is a valid Python and PyTorch subscript, so it resolves to the last page slot of the pool rather than raising — a transfer against another request’s KV. Two API points keep a page index from reaching device memory unchecked.
|
Yields |
|
The bytes of one page slot, raising |
Build transfer targets with valid_page_slots and address them with slot_tensor. This covers block_ids, cache_block_ids, RequestData.new_block_ids, and both *_by_layer_group forms.
from tensorrt_llm._torch.pyexecutor.connectors.kv_cache_layout import valid_page_slots
for ordinal, slot in valid_page_slots(cache_block_ids):
tokens = all_tokens[ordinal * tokens_per_block:(ordinal + 1) * tokens_per_block]
store.put(self._key(tokens), region.slot_tensor(slot))
Alignment is not the same as stability. Under a sliding window an entry that was reported with a page reads back as -1 once the window passes its block, and the delta — which carries only the ordinals appended since the last call — does not restate it. Save a block when its tokens complete rather than deferring: a completed block is far inside the window for any usable window size, whereas a deferred save can reach a slot the cache has already reclaimed.
What a connector can persist under a sliding window#
Under sliding-window attention, a connector can persist at most window_size tokens per sequence, not prompt_len.
The KV cache manager reclaims a block’s page once the window has moved past it, so by the time request_finished runs there is no readable KV for anything older than the last window_size tokens. Those ordinals report no page (see Blocks with no page), and the page slots offered to save from cover the live window only. A prefix-caching connector on such a model therefore caches a tail rather than a prefix, and the prefix it can serve back on a later request is bounded the same way.
This is a property of the cache, not of the connector: the blocks are gone whether or not a connector is attached. The same bound applies to the KV cache transceiver, which drops the same range before sending.
Serving a prefix#
A connector that implements only the flat callbacks and register_kv_caches runs unchanged on any model with a single non-sliding attention window. Where the cache describes itself as pools rather than one tensor it calls register_kv_cache_layout instead — but that method’s default reconstructs the single-pool tensor, in the same [num_blocks, num_layers, kv_factor, block_size] shape and KV dtype, and forwards it to register_kv_caches. The same applies to the two block-id callbacks: their per-layer-group forms default to the flat ones when there is a single layer group.
Variable sliding-window attention is the case where that stops working, because the cache then allocates one pool per window size and a page index is scoped to a layer group. See Running under VSWA.
A model whose layers all share one sliding window stays a single layer group, so the flat callbacks still apply and such a connector is not refused. What differs is that the callbacks cover the live window only. Blocks the window has passed report -1 (BAD_PAGE_INDEX) in place — the list stays aligned to block ordinals, so entry i still describes tokens [i * tokens_per_block, (i+1) * tokens_per_block), but the up-front blocks carry no page and are not available to load into or save from. Filter with valid_page_slots, described in Blocks with no page; without it a -1 resolves to the last page slot of the pool. A warning naming the window size is logged at start-up when a flat-only connector is attached to such a model.
get_num_new_matched_tokens is asked once the batch for the upcoming forward pass is final. A request that is asked is therefore a request that runs, and the connector can take ownership of remote blocks in the query and release it in request_finished.
Two things are worth knowing when tuning a deployment.
The runtime may honour less than you offer. With chunked prefill the cache is allocated per context chunk, which is what bounds its memory, so an offer reaching past the current chunk requires the runtime to grow the allocation and that can fail under pressure. The runtime then serves the part it can cover and computes the rest locally. The amount actually served is what
RequestData.computed_positionreflects; the unserved remainder needs no action from the connector beyond its usualrequest_finishedcleanup.The query is not part of the scheduler’s budget. The scheduler sizes a request’s chunk as if the connector will serve nothing, so a served prefix reduces the work in the forward pass but does not free budget for another request in the same iteration.
Specify enable_block_reuse=True alongside the connector for any of this to run; see Block reuse alongside the connector.
get_num_new_matched_tokens is called at most once per KV allocation. This is the precise form of the “once per request” rule: if a request’s KV cache is destroyed and the request is replayed – which MAX_UTILIZATION does under memory pressure – the replay asks again, because the pages the first answer described are gone.
Deployment note. Under a connector, a workload that was token-bound becomes KV-bound: the connector removes forward-pass tokens but its prefix still occupies GPU pages. Lowering max_num_tokens to hand memory back to the KV pool is usually the right adjustment, the opposite of the guidance for a connector-free deployment.
2. Worker Interface (KvCacheConnectorWorker)#
These methods run on all workers (GPU processes) and interact with the actual GPU data.
register_kv_caches(self, kv_cache_tensor: torch.Tensor)Description: Called at initialization. Provides the worker with the GPU KV cache tensors.
Arguments:
kv_cache_tensoris the underlying storage tensor for the KV cache, shaped[num_blocks, num_layers, kv_factor, block_size]. Rowblock_idis that block’s KV for every layer. Dimension 1 is indexed by model layer in ascending order, so thelayer_idxpassed towait_for_layer_loadandsave_kv_layerindexes it directly — no mapping is supplied, and none is needed.
register_kv_cache_layout(self, layout: KvCacheLayout)Description: Called at initialization instead of
register_kv_cacheswhen the cache describes itself as pools rather than one tensor.KvCacheLayoutgives byte ranges per layer group:layout.groups[g].regions[r], where the data for page slotiis atregion.base + region.stride * iforregion.sizebytes, orregion.slot_tensor(i, dtype). Full attribute reference:KvCacheLayoutreference.Default: reconstructs the single-pool tensor and forwards it to
register_kv_caches, so a connector that does not override this needs no changes for any single-window model. It raises when the cache cannot be described as one tensor — several layer groups (VSWA), or several regions (block scales, layers of differing size).
start_load_kv(self, stream: torch.cuda.Stream)Description: Initiates the loading of KV blocks from the external source into the GPU memory.
Arguments:
streamis the CUDA stream where the forward pass is executed in.
wait_for_layer_load(self, layer_idx: int, stream: torch.cuda.Stream)Description: A synchronization point. Ensures that the KV cache for a specific layer is fully loaded before the model attempts to perform the forward pass on that layer.
save_kv_layer(self, layer_idx: int, stream: torch.cuda.Stream)Description: Triggers the saving of a specific layer’s KV cache.
wait_for_save(self, stream: torch.cuda.Stream)Description: A synchronization point to ensure all save operations are enqueued or completed.
get_finished(self, finished_gen_req_ids, started_loading_req_ids) -> tuple[list[int], list[int]]Description: Polled by the runtime to check the status of asynchronous operations.
Returns: Two lists of request IDs: those that have finished saving, and those that have finished loading.
Example Implementation#
The file examples/llm-api/llm_kv_cache_connector.py provides a reference implementation of a Persistent KV Cache.
Overview#
This example implements a file-system based KV cache.
Save: When a request finishes or needs to be swapped out, its KV blocks are saved to disk as
.ptfiles.Load: When a new request arrives with the same prompt prefix, the connector identifies the cached files and loads them back into GPU memory, skipping re-computation.
Implementation Details#
Metadata: The example defines a
PersistentKvCacheConnectorMetadatadataclass containing lists of(file_path, block_id)tuples for both loading and saving. This simple structure allows the Scheduler to tell the Worker exactly which file corresponds to which GPU block index.Hashing Strategy: The
PersistentKvCacheConnectorLeaderhashes the token sequence of a block to generate a unique filename (e.g.,hash_value.pt). This acts as the lookup key.Worker Logic:
start_load_kv: Iterates through the load list provided in the metadata, loads the.ptfile to CPU, and copies it to the specificblock_idin the GPU tensor.wait_for_save: Performs the reverse. It copies data from the GPUblock_idto CPU and saves it to disk usingtorch.save.
Limitations & Patterns#
This example illustrates the API mechanics but has several limitations that make it unsuitable for high-performance production use without modification:
Blocking I/O: The example uses
torch.loadandtorch.savesynchronously. In a real implementation, these should be offloaded to a background thread or asynchronous I/O handler to avoid stalling the GPU.Simplified Block Matching: The
get_num_new_matched_tokensimplementation in the example only matches full blocks. It does not handle partial cache hits.FileSystem Latency: Storing one file per block can create high filesystem overhead.
Usage#
To run the example:
python examples/llm-api/llm_kv_cache_connector.py <model_path>
The script demonstrates:
Generating text for a prompt (First run).
Destroying the LLM instance.
Creating a new LLM instance with the same connector config.
Generating text for the same prompt (Second run).
Asserting that the outputs match, proving the state was correctly restored from the disk cache.