SkipSoftmaxAttentionConfig#

class tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig(
*,
algorithm: Literal['skip_softmax'] = 'skip_softmax',
threshold_scale_factor: float | Dict[str, float] | None = None,
target_sparsity: float | Dict[str, float] | None = None,
uses_spcompress: bool = False,
)[source]#

Bases: BaseSparseAttentionConfig

Configuration for skip softmax attention.

field algorithm: Literal['skip_softmax'] = 'skip_softmax'#
field target_sparsity: float | Dict[str, float] | None = None#

Target sparsity for prefill and/or decode phases. Requires formula coefficients in the model’s config.json. Ignored if threshold_scale_factor is also set.

field threshold_scale_factor: float | Dict[str, float] | None = None#

The threshold scale factor for skip softmax attention.

field uses_spcompress: bool = False#

Whether to enable spcompress (context phase, SM107 only).

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

supports_backend(backend: str) → bool[source]#

Override if the sparse attention algorithm does not support a subset of the possible backends.

to_sparse_metadata_params(**kwargs)#

Lower user-facing config into SparseMetadataParams.

to_sparse_params(**kwargs)[source]#

Lower user-facing config into SparseParams.

validator validate_target_sparsity  »  target_sparsity[source]#