MultimodalConfig#
- class tensorrt_llm.llmapi.MultimodalConfig(
- *,
- encoder_cuda_graph: dict[str, MultimodalEncoderCudaGraphConfig] | None = None,
- encoder_side_stream_max_ahead: Annotated[int, Ge(ge=0)] = 0,
- encoder_cache_max_bytes: Annotated[int, Ge(ge=0)] = 134217728,
- encoder_scheduling_policy: MultimodalEncoderSchedulingPolicy = MultimodalEncoderSchedulingPolicy.DEFAULT,
- video_pruning_rate: Annotated[float | None, Ge(ge=0.0), Lt(lt=1.0)] = None,
Bases:
StrictBaseModelMultimodal model configuration.
- field encoder_cache_max_bytes: Annotated[int, Ge(ge=0)] = 134217728#
Maximum bytes for the opt-in multimodal encoder embedding cache; 0 disables it. String values such as ‘512MB’ and ‘1GiB’ use binary units. Inline encoding caches whole single-modality requests; item scheduling caches individual items. Compatible with side-stream prefetch; their memory limits are additive.
- Constraints:
ge = 0
- field encoder_cuda_graph: dict[str, MultimodalEncoderCudaGraphConfig] | None = None#
CUDA graph capture for multimodal encoders, keyed by modality name. This config is not applied automatically - each model must read model_config.multimodal_config.encoder_cuda_graph and implement capture + replay via MultimodalEncoderCudaGraphRunner (see tensorrt_llm/_torch/models/multimodal_encoder_graph.py).
- field encoder_scheduling_policy: MultimodalEncoderSchedulingPolicy = MultimodalEncoderSchedulingPolicy.DEFAULT#
MM encoder scheduling policy for models that support item-level encoder scheduling. DISABLED: legacy inline encode (item scheduling and its byte budget off). DEFAULT: item scheduling. EAGER: item scheduling that advances encoder work for active requests before LLM capacity filtering. Ignored for models that do not support item scheduling.
- field encoder_side_stream_max_ahead: Annotated[int, Ge(ge=0)] = 0#
Maximum number of pending multimodal requests whose encoder work can be prefetched on a side CUDA stream ahead of admission. 0 disables side-stream prefetch. Incompatible with encoder_cuda_graph because graph replay uses static buffers. Can be combined with encoder_cache_max_bytes; the two memory limits are additive.
- Constraints:
ge = 0
- field video_pruning_rate: float | None = None#
Pruning rate for video frames in multimodal models for Efficient Video Sampling (EVS). NOTE: this is currently only implemented in nemotron multimodal models. None (default) disables EVS, values in [0, 1) enable pruning.
- Constraints:
ge = 0.0
lt = 1.0
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.