DeepSeekSparseAttentionConfig#

class tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig(
*,
algorithm: Literal['dsa'] = 'dsa',
seq_len_threshold: int | None = None,
index_n_heads: int | None = None,
index_head_dim: int | None = None,
index_topk: int | None = None,
indexer_max_chunk_size: int | None = None,
skip_indexer_for_short_seqs: bool = True,
use_cute_dsl_topk: bool = False,
use_cute_dsl_paged_mqa_logits: bool = False,
q_split_threshold: int = 8192,
indexer_rope_interleave: bool = False,
enable_heuristic_topk: bool = False,
use_self_sampling_topk: bool = True,
use_gvr_emission: bool = False,
indexer_k_dtype: Literal['fp8', 'fp4'] = 'fp8',
index_share_for_mtp_iteration: bool | None = None,
)[source]#

Bases: SeqLenAwareSparseAttentionConfig

Configuration for DeepSeek Sparse Attention.

field algorithm: Literal['dsa'] = 'dsa'#
field enable_heuristic_topk: bool = False#

Whether to enable Guess-Verify-Refine (GVR) Top-K for the DSA decode indexer instead of the exact insertion/radix Top-K path. Currently supported for index_topk ∈ {512, 1024, 2048} on Blackwell (SM100+), with compress_ratio ∈ {1, 4} (DSv3.2 + DSv4 indexers). Falls back to the production insertion/radix Top-K path when prerequisites are not met. use_self_sampling_topk selects the GVR engine generation.

field index_head_dim: int | None = None#

The dimension of the indexer heads.

field index_n_heads: int | None = None#

The number of heads for the indexer.

field index_share_for_mtp_iteration: bool | None = None#

Reuse the indexer Top-K across MTP draft steps instead of recomputing it each step. Defaults to the model’s HF config value.

field index_topk: int | None = None#

The topk for the indexer.

field indexer_k_dtype: Literal['fp8', 'fp4'] = 'fp8'#

Data type used for the indexer K cache. fp8 stores one FP8 E4M3 byte per element with a per-128 float32 scale; fp4 packs two FP4 E2M1 codes per byte with a per-32 UE8M0 exponent, halving the per-token indexer K footprint (132 B to 68 B at index_head_dim=128). fp4 requires Blackwell+ (SM>=100) at runtime and index_head_dim=128.

field indexer_max_chunk_size: int | None = None#

The maximum chunk size for the indexer.

field indexer_rope_interleave: bool = False#

Whether to use interleaved RoPE layout for the indexer.

field q_split_threshold: int = 8192#

If number of packed tokens in prefill chunk exceeds this threshold, q tokens will be evenly distributed across ranks for indexer computation. If negative, q split will always be disabled.

field seq_len_threshold: int | None = None#

The sequence length threshold for separating short and long sequences.

field skip_indexer_for_short_seqs: bool = True#

Whether to skip the MQA and Top-K in the indexer for short sequences.

field use_cute_dsl_paged_mqa_logits: bool = False#

Whether to use CuTE DSL paged MQA logits kernel on SM100-family GPUs instead of C++ DeepGEMM.

field use_cute_dsl_topk: bool = False#

Whether to use CuTE DSL top-k kernel instead of the CUDA C++ indexer_topk_decode.

field use_gvr_emission: bool = False#

Enable the emission-assisted block-skip optimization for the temporal-hint GVR engine. When set, the FP4 indexer epilogue emits per-block max logits so the GVR Top-K can skip whole blocks. Only takes effect with enable_heuristic_topk=True, use_self_sampling_topk=False, and the FP4 paged-MQA-logits path; ignored otherwise. The self-sampling engine derives its bracket from the current row and does not use emission.

field use_self_sampling_topk: bool = True#

Select the GVR engine generation when enable_heuristic_topk is set: True (default) runs the hint-free self-sampling engine, which derives its search bracket from the current row and keeps no cross-step state; False runs the temporal-hint engines, which reuse the previous decode step’s Top-K indices as hints. Ignored when enable_heuristic_topk is False.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

get_indices_block_size() → int#
needs_separate_short_long_cuda_graphs() → bool[source]#

Whether to capture separate CUDA graphs for short and long sequences. Use seq_len_threshold to determine the threshold for separating short and long sequences.

supports_backend(backend: str) → bool[source]#

Override if the sparse attention algorithm does not support a subset of the possible backends.

to_sparse_metadata_params(**kwargs)[source]#

Lower user-facing config into SparseMetadataParams.

to_sparse_params(**kwargs)[source]#

Lower user-facing config into SparseParams.