TorchCompileConfig#

class tensorrt_llm.llmapi.TorchCompileConfig(
*,
enable_fullgraph: bool = True,
enable_inductor: bool = False,
enable_piecewise_cuda_graph: bool = False,
capture_num_tokens: List[Annotated[int, Gt(gt=0)]] | None = None,
enable_userbuffers: bool = True,
max_num_streams: Annotated[int, Gt(gt=0)] = 3,
)[source]#

Bases: StrictBaseModel

Configuration for torch.compile.

field capture_num_tokens: List[Annotated[int, Gt(gt=0)]] | None = None#

Deprecated. Use prefill_capture_num_tokens instead.

field enable_fullgraph: bool = True#

Enable full graph compilation in torch.compile.

field enable_inductor: bool = False#

Enable inductor backend in torch.compile.

field enable_piecewise_cuda_graph: bool = False#

Deprecated. Use prefill_cuda_graph_backend=’piecewise’ instead.

field enable_userbuffers: bool = True#

When torch compile is enabled, userbuffers is enabled by default.

field max_num_streams: Annotated[int, Gt(gt=0)] = 3#

The maximum number of CUDA streams to use for torch.compile.

Constraints:
  • gt = 0

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

validator validate_capture_num_tokens  »  capture_num_tokens[source]#