BlockReuseConfig#
- class tensorrt_llm.llmapi.BlockReuseConfig(
- *,
- policy: Literal['all_reusable', 'per_request', 'per_conversation'] = 'all_reusable',
- max_num_turns: Annotated[int, Gt(gt=0)] = 1,
Bases:
StrictBaseModelConfiguration for KV cache block reuse policies.
- field max_num_turns: Annotated[int, Gt(gt=0)] = 1#
Maximum number of completed conversation turns whose committed SWA-window blocks and Mamba stable-boundary state are retained by KV cache manager v2. Only used when policy is ‘per_conversation’.
- Constraints:
gt = 0
- field policy: Literal['all_reusable', 'per_request', 'per_conversation'] = 'all_reusable'#
KV cache manager v2 block reuse policy. ‘all_reusable’ commits reusable blocks after every context chunk; ‘per_request’ commits them only after the final context chunk; ‘per_conversation’ uses ‘per_request’ commits and retains committed SWA-window blocks and Mamba stable-boundary state for up to max_num_turns completed turns. Periodic Mamba state snapshots are disabled with ‘per_conversation’. All reusable blocks remain subject to normal cache eviction. Requests without conversation params use ‘per_request’ behavior. When ‘all_reusable’ and SWA scratch reuse are both enabled, only non-scratch blocks are committed for reuse.
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.