ColdPageQuantizationCompressionConfig#

class tensorrt_llm.llmapi.ColdPageQuantizationCompressionConfig(
*,
algorithm: Literal['quantization_for_cold_page'] = 'quantization_for_cold_page',
quant: Literal['nvfp4'] = 'nvfp4',
scale_checkpoint_path: Annotated[str | None, MinLen(min_length=1)] = None,
)[source]#

Bases: KvCacheCompressionConfig

Quantize Host and Disk KV pages without changing the active GPU cache.

field algorithm: Literal['quantization_for_cold_page'] = 'quantization_for_cold_page'#
field quant: Literal['nvfp4'] = 'nvfp4'#

Quantization format stored in the compressed cache tier.

field scale_checkpoint_path: str | None = None#

Optional local ModelOpt NVFP4 checkpoint directory supplying per-layer K/V global scales. Omit it to use identity global scales.

Constraints:
  • min_length = 1

__init__(
**data: Any,
) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

supports_block_reuse() → bool[source]#

Block reuse is unchanged because token identity is preserved.

supports_speculative_decoding() → bool[source]#

Target and draft KVCMs encode their own cold pages independently.

changes_physical_kv_length: ClassVar[bool] = False#

Whether physical and logical KV lengths can diverge.