QSASparseAttentionConfig#

class tensorrt_llm.llmapi.QSASparseAttentionConfig(
*,
algorithm: Literal['qsa'] = 'qsa',
seq_len_threshold: Annotated[int | None, Gt(gt=0)] = None,
enable_heuristic_topk: bool = False,
index_share_for_mtp_iteration: bool | None = None,
)[source]#

Bases: SeqLenAwareSparseAttentionConfig

Configuration for QSA compressed query-selected attention.

field algorithm: Literal['qsa'] = 'qsa'#

Select QSA compressed query-selected sparse attention.

field enable_heuristic_topk: bool = False#

Whether to enable the Guess-Verify-Refine (GVR) Top-K for the QSA indexer instead of the exact radix Top-K. QSA dispatches only the hint-free self-sampling engine, which requires Blackwell (SM100/103), the CUTLASS DSL, a compressed-group budget (indexer_budget / indexer_compress_ratio) in {512, 1024, 2048}, and an indexer_compress_ratio of 4. Falls back to the exact radix Top-K with a one-time warning when the prerequisites are not met.

field index_share_for_mtp_iteration: bool | None = None#

Whether the MTP draft loop reuses the indexer selection captured by its draft-extend pass instead of re-running the indexer on every draft decode step. The query advances by at most max_draft_len positions, so the captured ranking is the one the indexer would recompute; compressed groups that complete during the loop are appended at lookup. When omitted, the checkpoint config supplies the value, defaulting to off.

field seq_len_threshold: int | None = None#

The sequence length threshold separating dense and QSA attention. When omitted, it is resolved to token_topk after checkpoint geometry is loaded; an explicit value below token_topk is raised to token_topk because it cannot reduce attention work.

Constraints:
  • gt = 0

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

get_indices_block_size() → int[source]#

Expanded QSA selections address individual tokens, not cache blocks.

needs_separate_short_long_cuda_graphs() → bool[source]#

Capture distinct dense and sparse decode graph families.

supports_backend(backend: str) → bool[source]#

Override if the sparse attention algorithm does not support a subset of the possible backends.

to_sparse_metadata_params(
**kwargs: object,
) → QSASparseMetadataParams[source]#

Lower user-facing config into SparseMetadataParams.

to_sparse_params(
**kwargs: object,
) → QSASparseParams[source]#

Lower user-facing config into SparseParams.