ColdPageQuantizationCompressionConfig#
- class tensorrt_llm.llmapi.ColdPageQuantizationCompressionConfig(
- *,
- algorithm: Literal['quantization_for_cold_page'] = 'quantization_for_cold_page',
- quant: Literal['nvfp4'] = 'nvfp4',
- scale_checkpoint_path: Annotated[str | None, MinLen(min_length=1)] = None,
Bases:
KvCacheCompressionConfigQuantize Host and Disk KV pages without changing the active GPU cache.
- field algorithm: Literal['quantization_for_cold_page'] = 'quantization_for_cold_page'#
- field quant: Literal['nvfp4'] = 'nvfp4'#
Quantization format stored in the compressed cache tier.
- field scale_checkpoint_path: str | None = None#
Optional local ModelOpt NVFP4 checkpoint directory supplying per-layer K/V global scales. Omit it to use identity global scales.
- Constraints:
min_length = 1
- __init__(
- **data: Any,
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- supports_speculative_decoding() bool[source]#
Target and draft KVCMs encode their own cold pages independently.
- changes_physical_kv_length: ClassVar[bool] = False#
Whether physical and logical KV lengths can diverge.