SAEnhancerConfig#
- class tensorrt_llm.llmapi.SAEnhancerConfig(
- *,
- threshold: Annotated[int, Gt(gt=0)] = 4,
- enable_global_pool: bool = False,
Bases:
StrictBaseModelConfiguration for the Suffix Automaton (SA) draft enhancer.
Use this to combine SA pattern-matching drafting with another speculative decoding method (Eagle3, MTP, PARD). When provided as
sa_configon a decoding config, SA drafting is enabled and may override neural draft tokens when the suffix match length meets the threshold.For standalone SA speculative decoding (no neural drafter), use
SADecodingConfiginstead.- field enable_global_pool: bool = False#
When True, each request searches all active SA states for the longest match, not just its own. Improves acceptance rates when requests share common patterns.
- field threshold: Annotated[int, Gt(gt=0)] = 4#
Minimum suffix match length required for the SA output to override neural draft tokens.
- Constraints:
gt = 0
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.