KvCacheConfig#

class tensorrt_llm.llmapi.KvCacheConfig(
*,
enable_block_reuse: bool = True,
max_tokens: int | None = None,
max_attention_window: ~typing.Annotated[~typing.List[int] | None,
~annotated_types.MinLen(min_length=1)] = None,
sink_token_length: int | None = None,
free_gpu_memory_fraction: ~typing.Annotated[float | None,
~annotated_types.Ge(ge=0),
~annotated_types.Le(le=1)] = 0.9,
host_cache_size: int | None = None,
disk_cache_size: ~typing.Annotated[int,
~annotated_types.Ge(ge=0)] | None = None,
disk_cache_path: str | None = None,
cross_kv_cache_fraction: float | None = None,
secondary_offload_min_priority: int | None = None,
event_buffer_max_size: int = 0,
attention_dp_events_gather_period_ms: int = 5,
kv_events_config: ~tensorrt_llm.llmapi.llm_args.KVEventsConfig | None = None,
enable_partial_reuse: bool = True,
copy_on_partial_reuse: bool = True,
use_uvm: bool = False,
max_gpu_total_bytes: ~typing.Annotated[int,
~annotated_types.Ge(ge=0)] = 0,
iteration_stats_interval: ~typing.Annotated[int,
~annotated_types.Gt(gt=0)] = 1,
dtype: str = 'auto',
mamba_ssm_cache_dtype: ~typing.Literal['auto',
'float16',
'bfloat16',
'float32'] = 'auto',
mamba_ssm_stochastic_rounding: bool = False,
mamba_ssm_philox_rounds: ~typing.Annotated[int,
~annotated_types.Ge(ge=1)] = 10,
tokens_per_block: int = 32,
mamba_state_cache_interval: ~typing.Annotated[int,
~annotated_types.Ge(ge=0)] | None = None,
mamba_state_config: ~tensorrt_llm.llmapi.llm_args.MambaStateConfig = <factory>,
use_kv_cache_manager_v2: bool | ~typing.Literal['auto'] = 'auto',
enable_swa_scratch_reuse: bool = False,
kv_cache_event_hash_algo: ~typing.Literal['auto',
'v1_block_key',
'v2_sha256',
'v2_sha256_64'] = 'auto',
max_util_for_resume: ~typing.Annotated[float,
~annotated_types.Gt(gt=0),
~annotated_types.Le(le=1)] = 0.95,
enable_kv_pool_rebalance: bool = False,
disk_prefetch_num_reqs: ~typing.Annotated[int,
~annotated_types.Ge(ge=0)] = 0,
fp8_context_mla_kv_len_cap: int | None = None,
pool_ratio: ~typing.Annotated[~typing.List[float] | None,
~annotated_types.MinLen(min_length=1)] = None,
avg_seq_len: ~typing.Annotated[int,
~annotated_types.Gt(gt=0)] | None = None,
block_reuse_config: ~tensorrt_llm.llmapi.llm_args.BlockReuseConfig = <factory>,
)[source]#

Bases: StrictBaseModel, PybindMirror

Configuration for the KV cache.

field attention_dp_events_gather_period_ms: int = 5#

The period in milliseconds to gather attention DP events across ranks.

field avg_seq_len: Annotated[int, Gt(gt=0)] | None = None#

Average total sequence length of the serving workload, used to build the KV cache manager v2 typical step for hybrid Mamba models and DeepSeek-V4. Hybrid Mamba models warn and fall back to half of max_seq_len when this is unset. This does not take effect when pool_ratio is set.

field block_reuse_config: BlockReuseConfig [Optional]#

KV cache manager v2 configuration for block reuse policies.

field copy_on_partial_reuse: bool = True#

Whether partially matched blocks that are in use can be reused after copying them.

field cross_kv_cache_fraction: float | None = None#

The fraction of the KV Cache memory should be reserved for cross attention. If set to p, self attention will use 1-p of KV Cache memory and cross attention will use p of KV Cache memory. Defaults to None (unset); must be set when using an encoder-decoder model and must not be set otherwise.

field disk_cache_path: str | None = None#

Directory used for disk KV cache files. Must be set when disk_cache_size is positive.

field disk_cache_size: Annotated[int, Ge(ge=0)] | None = None#

Size of the disk cache in bytes. Only used by KV cache manager v2 in the PyTorch backend.

field disk_prefetch_num_reqs: int = 0#

Number of queued context requests to prefetch disk-tier KV cache blocks to host for. Set to 0 to disable prefetch. Only effective with KV cache manager v2 and block reuse enabled.

Constraints:
  • ge = 0

field dtype: str = 'auto'#

The data type for the KV cache. ‘auto’ (default) leaves the checkpoint’s own KV-cache quantization metadata untouched (quant_config.kv_cache_quant_algo is inherited as-is); ‘fp8’, ‘fp8_ds_mla’, or ‘nvfp4’ override it explicitly. ‘fp8_ds_mla’ selects the packed FP8 cache used by sparse MLA on SM90/SM120/SM121. Resolved at LLM-construction time, including when set via trtllm-serve –extra_llm_api_options.

field enable_block_reuse: bool = True#

Controls if KV cache blocks can be reused for different requests.

field enable_kv_pool_rebalance: bool = False#

Opt in to the KVCacheManagerV2 auto-tuner (adjust()) for rebalancing pool-group ratios between iterations. When True the PyExecutor calls adjust() opportunistically; the auto-tuner itself remains gated by V2’s internal 2000-sample / 120s cooldown. When False (default) the rebalance hook is skipped entirely and pool ratios remain at their warmup-derived values. Beta: enable at your own risk. Only used when using KV cache manager v2 (experimental). This option is incompatible with dtype=’fp8_ds_mla’.

field enable_partial_reuse: bool = True#

Whether blocks that are only partially matched can be reused.

field enable_swa_scratch_reuse: bool = False#

Whether KV cache manager v2 uses SWA scratch reuse during prefill.

field event_buffer_max_size: int = 0#

Maximum size of the event buffer. If set to 0, the event buffer will not be used.

field fp8_context_mla_kv_len_cap: int | None = None#

Override, in tokens, for the max summed attended-KV length (total_kv_len) per forward step that the fp8 context-MLA attention workspace is reserved and scheduled for. Only affects fp8 context-MLA models (e.g. DeepSeek / Kimi with an fp8 KV cache). None (default) reserves for the never-stall worst case min(max_batch_size, max_num_tokens) * max_seq_len. A smaller value reserves less workspace (freeing KV cache) and defers context requests whose summed attended KV would exceed it; it is floored at max_seq_len and capped at the worst case. Safe at any value – the scheduler enforces it – trading prefill batching under heavy reuse for KV cache capacity.

field free_gpu_memory_fraction: float | None = 0.9#

The fraction of GPU memory fraction that should be allocated for the KV cache. Default is 90%. If both max_tokens and free_gpu_memory_fraction are specified, memory corresponding to the minimum will be used.

Constraints:
  • ge = 0

  • le = 1

field host_cache_size: int | None = None#

Size of the host cache in bytes. If both max_tokens and host_cache_size are specified, memory corresponding to the minimum will be used.

field iteration_stats_interval: Annotated[int, Gt(gt=0)] = 1#

How often (in iterations) to collect per-iteration KV cache statistics. A value of 1 means every iteration; a value of N means every Nth iteration. Between collections, the C++ deltas accumulate, so the reported deltas cover N iterations.

Constraints:
  • gt = 0

field kv_cache_event_hash_algo: Literal['auto', 'v1_block_key', 'v2_sha256', 'v2_sha256_64'] = 'auto'#

The block hash algorithm used by KV cache manager events. ‘auto’ uses the native hash for each KV cache manager. Explicit V2 hash choices are ignored with a warning by the V1 KV cache manager.

field kv_events_config: KVEventsConfig | None = None#

Streaming (push-based) KV cache event publishing (KV cache manager V2 only). When set, each rank publishes its own events directly (e.g. over ZeroMQ) instead of the buffered event_buffer_max_size gather/poll path.

field mamba_ssm_cache_dtype: Literal['auto', 'float16', 'bfloat16', 'float32'] = 'auto'#

The data type to use for the Mamba SSM cache. If set to ‘auto’, the data type will be inferred from the model config.

field mamba_ssm_philox_rounds: int = 10#

Number of Philox rounds for stochastic rounding PRNG. Higher values give better randomness but increase compute cost. Only used when mamba_ssm_stochastic_rounding is enabled.

Constraints:
  • ge = 1

field mamba_ssm_stochastic_rounding: bool = False#

Enable stochastic rounding for Mamba SSM state updates. Only applicable with float16 cache dtype.

field mamba_state_cache_interval: Annotated[int, Ge(ge=0)] | None = None#

Deprecated alias for mamba_state_config.periodic_snapshot_interval.

field mamba_state_config: MambaStateConfig [Optional]#

Configuration for reusable Mamba state snapshots.

field max_attention_window: List[int] | None = None#

Size of the attention window for each sequence. Only the last tokens will be stored in the KV cache. If the number of elements in max_attention_window is less than the number of layers, max_attention_window will be repeated multiple times to the number of layers.

Constraints:
  • min_length = 1

field max_gpu_total_bytes: Annotated[int, Ge(ge=0)] = 0#

The maximum size in bytes of GPU memory that can be allocated for the KV cache. If both max_gpu_total_bytes and free_gpu_memory_fraction are specified, memory corresponding to the minimum will be allocated.

Constraints:
  • ge = 0

field max_tokens: int | None = None#

The maximum number of tokens that should be stored in the KV cache. If both max_tokens and free_gpu_memory_fraction are specified, memory corresponding to the minimum will be used.

field max_util_for_resume: float = 0.95#

The maximum utilization of the KV cache for resume. Default is 95%. Only used when using KV cache manager v2 (experimental).

Constraints:
  • gt = 0

  • le = 1

field pool_ratio: List[float] | None = None#

Initial hot-tier byte ratios by layer group for KV cache manager v2. Values map to KVCacheManagerV2 layer-group ID order and must sum to 1.0. Cold tiers preserve the implied slot-count ratios. Hybrid Mamba models and DeepSeek-V4 use this directly, so avg_seq_len does not take effect when this is set.

Constraints:
  • min_length = 1

field secondary_offload_min_priority: int | None = None#

Only blocks with priority > secondary_offload_min_priority can be offloaded to secondary memory.

field tokens_per_block: int = 32#

The number of tokens per block.

field use_kv_cache_manager_v2: bool | Literal['auto'] = 'auto'#

Whether to use the KV cache manager v2 (experimental). ‘auto’ uses the model-specific preference and falls back to False when the model does not declare one.

field use_uvm: bool = False#

Whether to use UVM for the KV cache.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

validator disable_periodic_mamba_snapshots_for_conversations  »  all fields[source]#

Use only explicit stable boundaries for conversation reuse.

classmethod from_pybind(
pybind_instance: PybindMirror,
) → T#

Construct an instance of the given class from the fields in the given pybind class instance.

Parameters:
  • cls – Type of the class to construct, must be a subclass of pydantic BaseModel

  • pybind_instance – Instance of the pybind class to construct from its fields

Notes

When a field value is None in the pybind class, but it’s not optional and has a default value in the BaseModel class, it would get the default value defined in the BaseModel class.

Returns:

Instance of the given class, populated with the fields of the given pybind instance

static get_pybind_enum_fields(pybind_class)#

Get all the enum fields from the pybind class.

static get_pybind_variable_fields(config_cls)#

Get all the variable fields from the pybind class.

static maybe_to_pybind(ins)#
validator migrate_legacy_mamba_interval  »  all fields[source]#

Copy the deprecated Mamba interval into its nested replacement.

static mirror_pybind_enum(pybind_class)#

Mirror the enum fields from the pybind class to the Python class.

static mirror_pybind_fields(pybind_class)#

Class decorator that ensures Python class fields mirror those of a C++ class.

Parameters:

pybind_class – The C++ class whose fields should be mirrored

Returns:

A decorator function that validates field mirroring

static pybind_equals(obj0, obj1)#

Check if two pybind objects are equal.

validator reject_fp8_ds_mla_pool_rebalance  »  all fields[source]#

Reject resizing while packed sparse-MLA pool views are fixed.

validator validate_cross_kv_cache_fraction  »  cross_kv_cache_fraction[source]#
validator validate_disk_cache_config  »  all fields[source]#
validator validate_dtype  »  dtype[source]#
validator validate_free_gpu_memory_fraction  »  free_gpu_memory_fraction[source]#

Validates that the fraction is between 0.0 and 1.0.

validator validate_mamba_snapshot_offsets  »  all fields[source]#
validator validate_max_attention_window  »  max_attention_window[source]#
validator validate_max_gpu_total_bytes  »  max_gpu_total_bytes[source]#
validator validate_pool_ratio  »  pool_ratio[source]#
sink_token_length: int | None#

Read-only data descriptor used to emit a runtime deprecation warning before accessing a deprecated field.

msg#

The deprecation message to be emitted.

wrapped_property#

The property instance if the deprecated field is a computed field, or None.

field_name#

The name of the field being deprecated.