TriAttentionKvCacheCompressionConfig#

class tensorrt_llm.llmapi.TriAttentionKvCacheCompressionConfig(
*,
algorithm: Literal['triattention'] = 'triattention',
eviction_mode: Literal['union', 'per_head', 'per_layer_perhead'] = 'union',
normalize_scores: bool = True,
budget: Annotated[int, Gt(gt=0)] = 2048,
beta: Annotated[int, Gt(gt=0)] = 128,
calibration_path: Annotated[str, MinLen(min_length=1)],
)[source]#

Bases: KvCacheCompressionConfig

TriAttention KV-cache compression: periodic decode-time eviction.

Scored by offline calibration (github.com/WeianMao/triattention; supply the official .pt via calibration_path). Pure compression — decode runs the model’s standard attention over the compacted cache.

field algorithm: Literal['triattention'] = 'triattention'#
field beta: int = 128#

Eviction period in confirmed generation tokens (upstream divide_length): one speculative iteration may advance the counter by multiple accepted tokens; at most one eviction is coalesced per update.

Constraints:
  • gt = 0

field budget: int = 2048#

Tokens kept at each periodic eviction; prompt tokens are always preserved on top.

Constraints:
  • gt = 0

field calibration_path: str [Required]#

Path to the official TriAttention calibration .pt (produced by github.com/WeianMao/triattention). TRT-LLM does not compute calibration; it converts this file to the runtime schema at load.

Constraints:
  • min_length = 1

field eviction_mode: Literal['union', 'per_head', 'per_layer_perhead'] = 'union'#

Which token set each eviction round keeps. union (default) takes the union of each KV head’s top-B and re-ranks it by the per-token max score; it matches the official base setting (per-head and per-layer-per-head pruning both off). per_head keeps a per-KV-head set shared across layers (mean of per-layer max); per_layer_perhead keeps a fully independent set per (layer, KV head).

field normalize_scores: bool = True#

Z-normalize each head’s scores over the decode region before selection (upstream default). union eviction requires True: its fused score+stats+union pipeline always normalizes.

__init__(
**data: Any,
) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

supports_block_reuse() → bool[source]#
supports_speculative_decoding() → bool[source]#
changes_physical_kv_length: ClassVar[bool] = True#

Whether physical and logical KV lengths can diverge.