BlockReuseConfig#

class tensorrt_llm.llmapi.BlockReuseConfig(
*,
policy: Literal['all_reusable', 'per_request', 'per_conversation'] = 'all_reusable',
max_num_turns: Annotated[int, Gt(gt=0)] = 1,
)[source]#

Bases: StrictBaseModel

Configuration for KV cache block reuse policies.

field max_num_turns: Annotated[int, Gt(gt=0)] = 1#

Maximum number of completed conversation turns whose committed SWA-window blocks and Mamba stable-boundary state are retained by KV cache manager v2. Only used when policy is ‘per_conversation’.

Constraints:
  • gt = 0

field policy: Literal['all_reusable', 'per_request', 'per_conversation'] = 'all_reusable'#

KV cache manager v2 block reuse policy. ‘all_reusable’ commits reusable blocks after every context chunk; ‘per_request’ commits them only after the final context chunk; ‘per_conversation’ uses ‘per_request’ commits and retains committed SWA-window blocks and Mamba stable-boundary state for up to max_num_turns completed turns. Periodic Mamba state snapshots are disabled with ‘per_conversation’. All reusable blocks remain subject to normal cache eviction. Requests without conversation params use ‘per_request’ behavior. When ‘all_reusable’ and SWA scratch reuse are both enabled, only non-scratch blocks are committed for reuse.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.