QSASparseAttentionConfig#

class tensorrt_llm.llmapi.QSASparseAttentionConfig(
*,
algorithm: Literal['qsa'] = 'qsa',
seq_len_threshold: Annotated[int | None, Gt(gt=0)] = None,
)[source]#

Bases: SeqLenAwareSparseAttentionConfig

Configuration for QSA compressed query-selected attention.

field algorithm: Literal['qsa'] = 'qsa'#

Select QSA compressed query-selected sparse attention.

field seq_len_threshold: int | None = None#

The sequence length threshold separating dense and QSA attention. When omitted, it is resolved to token_topk after checkpoint geometry is loaded; an explicit value below token_topk is raised to token_topk because it cannot reduce attention work.

Constraints:
  • gt = 0

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

get_indices_block_size() → int[source]#

Expanded QSA selections address individual tokens, not cache blocks.

needs_separate_short_long_cuda_graphs() → bool[source]#

Capture distinct dense and sparse decode graph families.

supports_backend(backend: str) → bool[source]#

Override if the sparse attention algorithm does not support a subset of the possible backends.

to_sparse_metadata_params(
**kwargs: object,
) → QSASparseMetadataParams[source]#

Lower user-facing config into SparseMetadataParams.

to_sparse_params(
**kwargs: object,
) → QSASparseParams[source]#

Lower user-facing config into SparseParams.