DeepSeekV4SparseAttentionConfig#

class tensorrt_llm.llmapi.DeepSeekV4SparseAttentionConfig(
*,
algorithm: ~typing.Literal['deepseek_v4'] = 'deepseek_v4',
seq_len_threshold: int | None = None,
index_n_heads: int | None = None,
index_head_dim: int | None = 128,
index_topk: int | None = 512,
indexer_max_chunk_size: int | None = None,
skip_indexer_for_short_seqs: bool = False,
use_cute_dsl_topk: bool = False,
use_cute_dsl_paged_mqa_logits: bool = False,
q_split_threshold: int = 8192,
indexer_rope_interleave: bool = False,
enable_heuristic_topk: bool = False,
use_self_sampling_topk: bool = True,
use_gvr_emission: bool = False,
indexer_k_dtype: ~typing.Literal['fp8',
'fp4'] = 'fp4',
index_share_for_mtp_iteration: bool | None = None,
compress_ratios: ~typing.List[int] = <factory>,
window_size: int = 128,
)[source]#

Bases: DeepSeekSparseAttentionConfig

Configuration for DeepSeek-V4 Sparse Attention.

field algorithm: Literal['deepseek_v4'] = 'deepseek_v4'#
field compress_ratios: List[int] [Optional]#

The compress ratios of each layer. DeepSeek-V4 uses 0 for uncompressed/SWA-only layers; the LLM API config normalizes 0 to 1, while checkpoint-facing semantics remain unchanged.

field enable_heuristic_topk: bool = False#

Whether to enable Guess-Verify-Refine (GVR) Top-K for the DSA decode indexer instead of the exact insertion/radix Top-K path. Currently supported for index_topk ∈ {512, 1024, 2048} on Blackwell (SM100+), with compress_ratio ∈ {1, 4} (DSv3.2 + DSv4 indexers). Falls back to the production insertion/radix Top-K path when prerequisites are not met. use_self_sampling_topk selects the GVR engine generation.

field index_head_dim: int | None = 128#

The dimension of the DeepSeek-V4 indexer heads.

field index_n_heads: int | None = None#

The number of heads for the indexer.

field index_share_for_mtp_iteration: bool | None = None#

Reuse the indexer Top-K across MTP draft steps instead of recomputing it each step. Defaults to the model’s HF config value.

field index_topk: int | None = 512#

The top-k for the indexer.

field indexer_k_dtype: Literal['fp8', 'fp4'] = 'fp4'#

Data type used for the indexer K cache. DeepSeek-V4 defaults to fp4 to reduce the per-token indexer K footprint on Blackwell+ (SM>=100). Set to fp8 for the legacy FP8 indexer K cache path.

field indexer_max_chunk_size: int | None = None#

The maximum chunk size for the indexer.

field indexer_rope_interleave: bool = False#

Whether to use interleaved RoPE layout for the indexer.

field q_split_threshold: int = 8192#

If number of packed tokens in prefill chunk exceeds this threshold, q tokens will be evenly distributed across ranks for indexer computation. If negative, q split will always be disabled.

field seq_len_threshold: int | None = None#

The sequence length threshold for separating short and long sequences.

field skip_indexer_for_short_seqs: bool = False#

Whether to skip the MQA and Top-K in the indexer for short sequences.

field use_cute_dsl_paged_mqa_logits: bool = False#

Whether to use CuTE DSL paged MQA logits kernel on SM100-family GPUs instead of C++ DeepGEMM.

field use_cute_dsl_topk: bool = False#

Whether to use CuTE DSL top-k kernel instead of the CUDA C++ indexer_topk_decode.

field use_gvr_emission: bool = False#

Enable the emission-assisted block-skip optimization for the temporal-hint GVR engine. When set, the FP4 indexer epilogue emits per-block max logits so the GVR Top-K can skip whole blocks. Only takes effect with enable_heuristic_topk=True, use_self_sampling_topk=False, and the FP4 paged-MQA-logits path; ignored otherwise. The self-sampling engine derives its bracket from the current row and does not use emission.

field use_self_sampling_topk: bool = True#

Select the GVR engine generation when enable_heuristic_topk is set: True (default) runs the hint-free self-sampling engine, which derives its search bracket from the current row and keeps no cross-step state; False runs the temporal-hint engines, which reuse the previous decode step’s Top-K indices as hints. Ignored when enable_heuristic_topk is False.

field window_size: int = 128#

The sliding window size in tokens for SWA layers.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

get_indices_block_size() → int#
needs_separate_short_long_cuda_graphs() → bool[source]#

Whether to capture separate CUDA graphs for short and long sequences. Use seq_len_threshold to determine the threshold for separating short and long sequences.

validator normalize_compress_ratios  »  compress_ratios[source]#
supports_backend(backend: str) → bool[source]#

Override if the sparse attention algorithm does not support a subset of the possible backends.

to_sparse_metadata_params(**kwargs)[source]#

Lower user-facing config into SparseMetadataParams.

to_sparse_params(**kwargs)[source]#

Lower user-facing config into SparseParams.

validator validate_index_head_dim  »  index_head_dim[source]#