MoeConfig#

class tensorrt_llm.llmapi.MoeConfig(
*,
backend: Literal['AUTO', 'CUTLASS', 'CUTEDSL', 'CUTEDSL_FC12', 'TRTLLM', 'DEEPGEMM', 'DENSEGEMM', 'VANILLA', 'TRITON', 'MARLIN', 'MEGAMOE_DEEPGEMM', 'MEGAMOE_CUTEDSL'] = 'AUTO',
max_num_tokens: int | None = None,
load_balancer: object | str | None = None,
disable_finalize_fusion: bool = False,
use_low_precision_moe_combine: bool = False,
)[source]#

Bases: StrictBaseModel

Configuration for MoE.

field backend: Literal['AUTO', 'CUTLASS', 'CUTEDSL', 'CUTEDSL_FC12', 'TRTLLM', 'DEEPGEMM', 'DENSEGEMM', 'VANILLA', 'TRITON', 'MARLIN', 'MEGAMOE_DEEPGEMM', 'MEGAMOE_CUTEDSL'] = 'AUTO'#

MoE backend to use. AUTO selects default backend based on model. It currently doesn’t always give the best choice for all scenarios. The capabilities of auto selection will be improved in future releases.

field disable_finalize_fusion: bool = False#

Disable FC2+finalize kernel fusion in CUTLASS MoE backend. Setting this to True recovers deterministic numerical behavior with top-k > 2.

field load_balancer: object | str | None = None#

Configuration for MoE load balancing.

field max_num_tokens: int | None = None#

If set, at most max_num_tokens tokens will be sent to torch.ops.trtllm.fused_moe at the same time. If the number of tokens exceeds max_num_tokens, the input tensors will be split into chunks and a for loop will be used.

field use_low_precision_moe_combine: bool = False#

Use low precision combine in MoE operations (only for NVFP4 quantization). When enabled, uses lower precision for combining expert outputs to improve performance.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.