DeepSeekV4SparseAttentionConfig#
- class tensorrt_llm.llmapi.DeepSeekV4SparseAttentionConfig(
- *,
- algorithm: ~typing.Literal['deepseek_v4'] = 'deepseek_v4',
- seq_len_threshold: int | None = None,
- index_n_heads: int | None = None,
- index_head_dim: int | None = 128,
- index_topk: int | None = 512,
- indexer_max_chunk_size: int | None = None,
- skip_indexer_for_short_seqs: bool = False,
- use_cute_dsl_topk: bool = False,
- use_cute_dsl_paged_mqa_logits: bool = False,
- q_split_threshold: int = 8192,
- indexer_rope_interleave: bool = False,
- enable_heuristic_topk: bool = False,
- use_self_sampling_topk: bool = True,
- use_gvr_emission: bool = False,
- indexer_k_dtype: ~typing.Literal['fp8',
- 'fp4'] = 'fp4',
- index_share_for_mtp_iteration: bool | None = None,
- compress_ratios: ~typing.List[int] = <factory>,
- window_size: int = 128,
Bases:
DeepSeekSparseAttentionConfigConfiguration for DeepSeek-V4 Sparse Attention.
- field algorithm: Literal['deepseek_v4'] = 'deepseek_v4'#
- field compress_ratios: List[int] [Optional]#
The compress ratios of each layer. DeepSeek-V4 uses 0 for uncompressed/SWA-only layers; the LLM API config normalizes 0 to 1, while checkpoint-facing semantics remain unchanged.
- field enable_heuristic_topk: bool = False#
Whether to enable Guess-Verify-Refine (GVR) Top-K for the DSA decode indexer instead of the exact insertion/radix Top-K path. Currently supported for index_topk ∈ {512, 1024, 2048} on Blackwell (SM100+), with compress_ratio ∈ {1, 4} (DSv3.2 + DSv4 indexers). Falls back to the production insertion/radix Top-K path when prerequisites are not met. use_self_sampling_topk selects the GVR engine generation.
- field index_head_dim: int | None = 128#
The dimension of the DeepSeek-V4 indexer heads.
- field index_n_heads: int | None = None#
The number of heads for the indexer.
Reuse the indexer Top-K across MTP draft steps instead of recomputing it each step. Defaults to the model’s HF config value.
- field index_topk: int | None = 512#
The top-k for the indexer.
- field indexer_k_dtype: Literal['fp8', 'fp4'] = 'fp4'#
Data type used for the indexer K cache. DeepSeek-V4 defaults to fp4 to reduce the per-token indexer K footprint on Blackwell+ (SM>=100). Set to fp8 for the legacy FP8 indexer K cache path.
- field indexer_max_chunk_size: int | None = None#
The maximum chunk size for the indexer.
- field indexer_rope_interleave: bool = False#
Whether to use interleaved RoPE layout for the indexer.
- field q_split_threshold: int = 8192#
If number of packed tokens in prefill chunk exceeds this threshold, q tokens will be evenly distributed across ranks for indexer computation. If negative, q split will always be disabled.
- field seq_len_threshold: int | None = None#
The sequence length threshold for separating short and long sequences.
- field skip_indexer_for_short_seqs: bool = False#
Whether to skip the MQA and Top-K in the indexer for short sequences.
- field use_cute_dsl_paged_mqa_logits: bool = False#
Whether to use CuTE DSL paged MQA logits kernel on SM100-family GPUs instead of C++ DeepGEMM.
- field use_cute_dsl_topk: bool = False#
Whether to use CuTE DSL top-k kernel instead of the CUDA C++ indexer_topk_decode.
- field use_gvr_emission: bool = False#
Enable the emission-assisted block-skip optimization for the temporal-hint GVR engine. When set, the FP4 indexer epilogue emits per-block max logits so the GVR Top-K can skip whole blocks. Only takes effect with enable_heuristic_topk=True, use_self_sampling_topk=False, and the FP4 paged-MQA-logits path; ignored otherwise. The self-sampling engine derives its bracket from the current row and does not use emission.
- field use_self_sampling_topk: bool = True#
Select the GVR engine generation when enable_heuristic_topk is set: True (default) runs the hint-free self-sampling engine, which derives its search bracket from the current row and keeps no cross-step state; False runs the temporal-hint engines, which reuse the previous decode step’s Top-K indices as hints. Ignored when enable_heuristic_topk is False.
- field window_size: int = 128#
The sliding window size in tokens for SWA layers.
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- get_indices_block_size() int#
- needs_separate_short_long_cuda_graphs() bool[source]#
Whether to capture separate CUDA graphs for short and long sequences. Use seq_len_threshold to determine the threshold for separating short and long sequences.