DecodeCudaGraphConfig#

class tensorrt_llm.llmapi.DecodeCudaGraphConfig(
*,
batch_sizes: List[int] | None = None,
max_batch_size: Annotated[int, Ge(ge=0)] = 0,
enable_padding: bool = False,
mode: Literal['decode'] = 'decode',
)[source]#

Bases: BaseCudaGraphConfig

CUDA graph configuration for decode requests.

field batch_sizes: List[int] | None = None#

List of batch sizes to create CUDA graphs for.

field enable_padding: bool = False#

If true, batches are rounded up to the nearest cuda_graph_batch_size. This is usually a net win for performance.

field max_batch_size: NonNegativeInt = 0#

Maximum batch size for CUDA graphs.

Constraints:
  • ge = 0

field mode: Literal['decode'] = 'decode'#

CUDA graph configuration mode.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

validator validate_base_cuda_graph_config  »  all fields#

Validate CUDA graph configuration.

Ensures that: 1. If batch_sizes is provided, max_batch_size is derived as max(batch_sizes).

If max_batch_size was already set it must be compatible (equal to max(batch_sizes)); otherwise an error is raised.

  1. If only max_batch_size is provided, batch_sizes is generated from it.

  2. If neither is provided, a default max_batch_size of 128 is used.