CacheTransceiverConfig#
- class tensorrt_llm.llmapi.CacheTransceiverConfig(
- *,
- backend: Literal['DEFAULT', 'UCX', 'NIXL', 'MOONCAKE', 'MPI'] | None = None,
- transceiver_runtime: Literal['CPP', 'PYTHON', 'auto'] | None = 'auto',
- max_tokens_in_buffer: int | None = None,
- kv_transfer_timeout_ms: Annotated[int, Gt(gt=0)] | None = 60000,
- kv_transfer_sender_future_timeout_ms: Annotated[int, Gt(gt=0)] | None = 1000,
- kv_transfer_poll_interval_ms: Annotated[int, Gt(gt=0)] | None = 5000,
- kv_cache_bounce_size_mb: Annotated[int, Ge(ge=0)] = 0,
- enable_pipelined_transfer: bool = False,
Bases:
StrictBaseModel,PybindMirrorConfiguration for the cache transceiver.
- field backend: Literal['DEFAULT', 'UCX', 'NIXL', 'MOONCAKE', 'MPI'] | None = None#
The communication backend type to use for the cache transceiver.
- field enable_pipelined_transfer: bool = False#
Transfer each completed prefill chunk’s KV cache while later chunks compute. Requires Python NIXL, generation-first scheduling, chunked prefill, pipeline_parallel_size=1, context_parallel_size=1 on both peers, beam_width=1, no bounce buffer or Mamba/hybrid cache, and block reuse disabled or set to all_reusable. Invalid static settings fail at startup; per-request constraints reject the request.
- field kv_cache_bounce_size_mb: int = 0#
Per-region size in MiB of the native-disagg KV-cache bounce buffer (one for send, one for recv). Bounce coalesces a request’s scattered per-block KV into one contiguous fabric-VMM buffer and issues a single multi-rail NIXL write. The size doubles as the on/off switch: 0 (default) keeps the per-block path, >0 enables bounce at that capacity. Only used by the Python (v2) transceiver.
- Constraints:
ge = 0
- field kv_transfer_poll_interval_ms: Annotated[int, Gt(gt=0)] | None = 5000#
Bounded wait interval in milliseconds for polling KV transfer progress when active transfers block disaggregated admission.
- field kv_transfer_sender_future_timeout_ms: Annotated[int, Gt(gt=0)] | None = 1000#
Duration in milliseconds of each bounded sender future wait slice while polling KV transfer completion. It does not set the overall transfer deadline.
- field kv_transfer_timeout_ms: Annotated[int, Gt(gt=0)] | None = 60000#
KV cache transfer timeout in milliseconds. Blocking sender waits use it as an absolute deadline; blocking receive task waits use it per task. The Python V2 transceiver requires a finite value; None remains available to other runtimes. It is distinct from the sender future wait slice.
- field max_tokens_in_buffer: int | None = None#
The max number of tokens the transfer buffer can fit.
- field transceiver_runtime: Literal['CPP', 'PYTHON', 'auto'] | None = 'auto'#
The runtime implementation. ‘auto’ (default) adopts the model’s preferred runtime when it declares one; otherwise it selects the Python transceiver, falling back to the C++ transceiver only when this config itself rules it out (non-NIXL backend or a null kv_transfer_timeout_ms) — any other incompatibility fails at transceiver creation. The fallback is decided independently on each server and is only logged, not surfaced, so keep context and generation server configurations consistent. ‘CPP’ selects the C++ transceiver, ‘PYTHON’ the Python transceiver. None is equivalent to ‘CPP’. ‘auto’ is only resolved on the PyTorch backend’s standard model-loading path; other paths (e.g. AutoDeploy) fall back to the C++ transceiver.
- __init__(**data: Any) None#
Create a new model by parsing and validating input data from keyword arguments.
Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.
self is explicitly positional-only to allow self as a field name.
- classmethod from_pybind(
- pybind_instance: PybindMirror,
Construct an instance of the given class from the fields in the given pybind class instance.
- Parameters:
cls – Type of the class to construct, must be a subclass of pydantic BaseModel
pybind_instance – Instance of the pybind class to construct from its fields
Notes
When a field value is None in the pybind class, but it’s not optional and has a default value in the BaseModel class, it would get the default value defined in the BaseModel class.
- Returns:
Instance of the given class, populated with the fields of the given pybind instance
- static get_pybind_enum_fields(pybind_class)#
Get all the enum fields from the pybind class.
- static get_pybind_variable_fields(config_cls)#
Get all the variable fields from the pybind class.
- static maybe_to_pybind(ins)#
- static mirror_pybind_enum(pybind_class)#
Mirror the enum fields from the pybind class to the Python class.
- static mirror_pybind_fields(pybind_class)#
Class decorator that ensures Python class fields mirror those of a C++ class.
- Parameters:
pybind_class – The C++ class whose fields should be mirrored
- Returns:
A decorator function that validates field mirroring
- static pybind_equals(obj0, obj1)#
Check if two pybind objects are equal.