TorchCompileConfig#
- class tensorrt_llm.llmapi.TorchCompileConfig(
- *,
- enable_fullgraph: bool = True,
- enable_inductor: bool = False,
- enable_piecewise_cuda_graph: bool = False,
- capture_num_tokens: List[Annotated[int, Gt(gt=0)]] | None = None,
- enable_userbuffers: bool = True,
- max_num_streams: Annotated[int, Gt(gt=0)] = 3,
Bases:
StrictBaseModelConfiguration for torch.compile.
- field capture_num_tokens: List[Annotated[int, Gt(gt=0)]] | None = None#
Deprecated. Use prefill_capture_num_tokens instead.
- field enable_fullgraph: bool = True#
Enable full graph compilation in torch.compile.
- field enable_inductor: bool = False#
Enable inductor backend in torch.compile.
- field enable_piecewise_cuda_graph: bool = False#
Deprecated. Use prefill_cuda_graph_backend=’piecewise’ instead.
- field enable_userbuffers: bool = True#
When torch compile is enabled, userbuffers is enabled by default.
- field max_num_streams: Annotated[int, Gt(gt=0)] = 3#
The maximum number of CUDA streams to use for torch.compile.
- Constraints:
gt = 0
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.