QSASparseAttentionConfig#
- class tensorrt_llm.llmapi.QSASparseAttentionConfig(
- *,
- algorithm: Literal['qsa'] = 'qsa',
- seq_len_threshold: Annotated[int | None, Gt(gt=0)] = None,
Bases:
SeqLenAwareSparseAttentionConfigConfiguration for QSA compressed query-selected attention.
- field algorithm: Literal['qsa'] = 'qsa'#
Select QSA compressed query-selected sparse attention.
- field seq_len_threshold: int | None = None#
The sequence length threshold separating dense and QSA attention. When omitted, it is resolved to token_topk after checkpoint geometry is loaded; an explicit value below token_topk is raised to token_topk because it cannot reduce attention work.
- Constraints:
gt = 0
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- get_indices_block_size() int[source]#
Expanded QSA selections address individual tokens, not cache blocks.
- needs_separate_short_long_cuda_graphs() bool[source]#
Capture distinct dense and sparse decode graph families.
- supports_backend(backend: str) bool[source]#
Override if the sparse attention algorithm does not support a subset of the possible backends.