SkipSoftmaxAttentionConfig#
- class tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig(
- *,
- algorithm: Literal['skip_softmax'] = 'skip_softmax',
- threshold_scale_factor: float | Dict[str, float] | None = None,
- target_sparsity: float | Dict[str, float] | None = None,
- uses_spcompress: bool = False,
Bases:
BaseSparseAttentionConfigConfiguration for skip softmax attention.
- field algorithm: Literal['skip_softmax'] = 'skip_softmax'#
- field target_sparsity: float | Dict[str, float] | None = None#
Target sparsity for prefill and/or decode phases. Requires formula coefficients in the model’s config.json. Ignored if threshold_scale_factor is also set.
- field threshold_scale_factor: float | Dict[str, float] | None = None#
The threshold scale factor for skip softmax attention.
- field uses_spcompress: bool = False#
Whether to enable spcompress (context phase, SM107 only).
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- supports_backend(backend: str) bool[source]#
Override if the sparse attention algorithm does not support a subset of the possible backends.
- to_sparse_metadata_params(**kwargs)#
Lower user-facing config into SparseMetadataParams.