QSASparseAttentionConfig#
- class tensorrt_llm.llmapi.QSASparseAttentionConfig(
- *,
- algorithm: Literal['qsa'] = 'qsa',
- seq_len_threshold: Annotated[int | None, Gt(gt=0)] = None,
- enable_heuristic_topk: bool = False,
- index_share_for_mtp_iteration: bool | None = None,
Bases:
SeqLenAwareSparseAttentionConfigConfiguration for QSA compressed query-selected attention.
- field algorithm: Literal['qsa'] = 'qsa'#
Select QSA compressed query-selected sparse attention.
- field enable_heuristic_topk: bool = False#
Whether to enable the Guess-Verify-Refine (GVR) Top-K for the QSA indexer instead of the exact radix Top-K. QSA dispatches only the hint-free self-sampling engine, which requires Blackwell (SM100/103), the CUTLASS DSL, a compressed-group budget (indexer_budget / indexer_compress_ratio) in {512, 1024, 2048}, and an indexer_compress_ratio of 4. Falls back to the exact radix Top-K with a one-time warning when the prerequisites are not met.
Whether the MTP draft loop reuses the indexer selection captured by its draft-extend pass instead of re-running the indexer on every draft decode step. The query advances by at most max_draft_len positions, so the captured ranking is the one the indexer would recompute; compressed groups that complete during the loop are appended at lookup. When omitted, the checkpoint config supplies the value, defaulting to off.
- field seq_len_threshold: int | None = None#
The sequence length threshold separating dense and QSA attention. When omitted, it is resolved to token_topk after checkpoint geometry is loaded; an explicit value below token_topk is raised to token_topk because it cannot reduce attention work.
- Constraints:
gt = 0
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- get_indices_block_size() int[source]#
Expanded QSA selections address individual tokens, not cache blocks.
- needs_separate_short_long_cuda_graphs() bool[source]#
Capture distinct dense and sparse decode graph families.
- supports_backend(backend: str) bool[source]#
Override if the sparse attention algorithm does not support a subset of the possible backends.