AttentionDpConfig#

class tensorrt_llm.llmapi.AttentionDpConfig(
*,
enable_balance: bool = False,
timeout_iters: int = 50,
batching_wait_iters: int = 10,
enable_kv_cache_aware_routing: bool = False,
kv_cache_routing_load_balance_weight: float = 1.0,
kv_cache_routing_match_rate_threshold: float = 0.1,
kv_cache_routing_fair_share_multiplier: float = 2.0,
kv_cache_routing_cold_start_warmup: bool = False,
kv_cache_routing_account_for_in_transfer: bool = False,
kv_cache_routing_conversation_affinity: bool = False,
kv_cache_routing_max_sessions: int = 65536,
kv_cache_routing_new_conv_placement: Literal['round_robin', 'least_queued'] = 'round_robin',
)[source]#

Bases: StrictBaseModel

Configuration for attention DP.

field batching_wait_iters: int = 10#

The number of iterations to wait for batching.

field enable_balance: bool = False#

Whether to enable balance.

field enable_kv_cache_aware_routing: bool = False#

Enable internal KV cache-aware routing for attention DP. When enabled, distributes requests among ranks within a single instance’s attention DP group, routing them to the rank with the matching prefix KV cache to reduce redundant prefill computation.

field kv_cache_routing_account_for_in_transfer: bool = False#

In-transfer load accounting in KV cache-aware routing. When True, requests still streaming KV to the GEN worker (tracked by the PyExecutor AsyncTransferManager but no longer in active_requests) are folded back into the per-rank load reported via RankState. This can improve balance under heavy disagg traffic but inflates num_active_requests reported upstream, which lets the inference loop’s idle-fetch wait expire even when no requests are runnable and causes empty fetch cycles. Default False preserves the prior behaviour (fetch blocks when truly idle). Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_cold_start_warmup: bool = False#

Cold-start mitigation in KV cache-aware routing. When True, the first tp_size relaxed requests after router init are round-robined across ranks (bypassing cache-affinity scoring) so every rank caches the shared system prompt before scoring would otherwise pin all traffic to the first warm rank. Only useful when requests share a long system prefix; for diverse-prompt workloads this can scatter requests that would otherwise consolidate on a single warm rank, wasting prefill. Default False preserves pre-warmup routing. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_conversation_affinity: bool = False#

Enable explicit conversation-affinity routing for attention DP. When True, the first request of each conversation is round-robined across ranks and every subsequent request carrying the same conversation_params.conversation_id is pinned to that conversation’s first-turn rank. OpenAI requests use the body conversation_params as canonical; the serve edge only creates conversation_params from the X-Session-ID header when the body does not provide it. This keeps a multi-turn conversation’s KV-cache prefix on one rank (maximizing block reuse, minimizing cross-rank migration). Unlike enable_kv_cache_aware_routing (affinity inferred from prefix-match length, which is lost when blocks are evicted), the conversation->rank map is explicit and survives eviction. Falls back to load-balanced round-robin when no conversation_id is available. Takes precedence over enable_kv_cache_aware_routing when both are set.

field kv_cache_routing_fair_share_multiplier: float = 2.0#

Loose per-rank active-request cap in KV cache-aware routing, expressed as a multiplier of the ceil fair-share (ceil((total_active + new) / tp_size)). Once a rank hits this cap within a scheduling batch it is removed from the eligible set for the remainder of the batch. Default 2.0 permits a 2x slack so cache affinity can dominate while preventing runaway concentration on a single rank. Set to 1.0 for strict fair share. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_load_balance_weight: float = 1.0#

Weight (beta) for the load-balance term in KV cache-aware routing. Higher values prioritize load balance over cache affinity. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_match_rate_threshold: float = 0.1#

Cache-affinity gate in KV cache-aware routing. For each request, match_len contributes to scoring only when max(match_len) / request_tokens across eligible ranks is strictly above this threshold; otherwise match_len is forced to 0 so routing is driven purely by load. Default 0.1 requires at least a 10% hit rate before cache affinity kicks in, which prevents a small universal prefix (e.g. a shared system prompt) from pinning all traffic to the first warm ranks. Set to 0.0 to honour any nonzero match. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_max_sessions: int = 65536#

LRU cap on the conversation->rank map used by conversation-affinity routing. The oldest conversations are evicted once more than this many are tracked, bounding memory on long-running servers. Only used when kv_cache_routing_conversation_affinity is True.

field kv_cache_routing_new_conv_placement: Literal['round_robin', 'least_queued'] = 'round_robin'#

Placement policy in conversation-affinity routing for requests with no pinned rank yet (first turn of a conversation, requests without a conversation_id, sticky overflow). ‘round_robin’ (default) equalizes per-rank conversation counts. ‘least_queued’ places them on the rank with the fewest live requests instead: per-conversation load (turn rate, fan-out, prefill length) is not uniform, so count-uniform round-robin can leave some ranks with deep queues while others idle; steering new conversations by queue depth evens that out and cuts tail TTFT. Existing conversation->rank pins are unaffected. Only used when kv_cache_routing_conversation_affinity is True.

field timeout_iters: int = 50#

The number of iterations to timeout.

class Config#

Bases: object

extra = 'forbid'#
__init__(**data: Any) None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

classmethod construct(
_fields_set: set[str] | None = None,
**values: Any,
) Self#
copy(
*,
include: AbstractSetIntStr | MappingIntStrAny | None = None,
exclude: AbstractSetIntStr | MappingIntStrAny | None = None,
update: Dict[str, Any] | None = None,
deep: bool = False,
) Self#

Returns a copy of the model.

!!! warning “Deprecated”

This method is now deprecated; use model_copy instead.

If you need include or exclude, use:

`python {test="skip" lint="skip"} data = self.model_dump(include=include, exclude=exclude, round_trip=True) data = {**data, **(update or {})} copied = self.model_validate(data) `

Parameters:
  • include – Optional set or mapping specifying which fields to include in the copied model.

  • exclude – Optional set or mapping specifying which fields to exclude in the copied model.

  • update – Optional dictionary of field-value pairs to override field values in the copied model.

  • deep – If True, the values of fields that are Pydantic models will be deep-copied.

Returns:

A copy of the model with included, excluded and updated fields as specified.

dict(
*,
include: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
exclude: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
by_alias: bool = False,
exclude_unset: bool = False,
exclude_defaults: bool = False,
exclude_none: bool = False,
) Dict[str, Any]#
classmethod from_orm(obj: Any) Self#
json(
*,
include: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
exclude: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
by_alias: bool = False,
exclude_unset: bool = False,
exclude_defaults: bool = False,
exclude_none: bool = False,
encoder: Callable[[Any], Any] | None = PydanticUndefined,
models_as_dict: bool = PydanticUndefined,
**dumps_kwargs: Any,
) str#
classmethod model_construct(
_fields_set: set[str] | None = None,
**values: Any,
) Self#

Creates a new instance of the Model class with validated data.

Creates a new model setting __dict__ and __pydantic_fields_set__ from trusted or pre-validated data. Default values are respected, but no other validation is performed.

!!! note

model_construct() generally respects the model_config.extra setting on the provided model. That is, if model_config.extra == ‘allow’, then all extra passed values are added to the model instance’s __dict__ and __pydantic_extra__ fields. If model_config.extra == ‘ignore’ (the default), then all extra passed values are ignored. Because no validation is performed with a call to model_construct(), having model_config.extra == ‘forbid’ does not result in an error if extra values are passed, but they will be ignored.

Parameters:
  • _fields_set – A set of field names that were originally explicitly set during instantiation. If provided, this is directly used for the [model_fields_set][pydantic.BaseModel.model_fields_set] attribute. Otherwise, the field names from the values argument will be used.

  • values – Trusted or pre-validated data dictionary.

Returns:

A new instance of the Model class with validated data.

model_copy(
*,
update: Mapping[str, Any] | None = None,
deep: bool = False,
) Self#
!!! abstract “Usage Documentation”

[model_copy](../concepts/models.md#model-copy)

Returns a copy of the model.

!!! note

The underlying instance’s [__dict__][object.__dict__] attribute is copied. This might have unexpected side effects if you store anything in it, on top of the model fields (e.g. the value of [cached properties][functools.cached_property]).

Parameters:
  • update – Values to change/add in the new model. Note: the data is not validated before creating the new model. You should trust this data.

  • deep – Set to True to make a deep copy of the model.

Returns:

New model instance.

model_dump(
*,
mode: Literal['json', 'python'] | str = 'python',
include: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
exclude: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
context: Any | None = None,
by_alias: bool | None = None,
exclude_unset: bool = False,
exclude_defaults: bool = False,
exclude_none: bool = False,
exclude_computed_fields: bool = False,
round_trip: bool = False,
warnings: bool | Literal['none', 'warn', 'error'] = True,
fallback: Callable[[Any], Any] | None = None,
serialize_as_any: bool = False,
polymorphic_serialization: bool | None = None,
) dict[str, Any]#
!!! abstract “Usage Documentation”

[model_dump](../concepts/serialization.md#python-mode)

Generate a dictionary representation of the model, optionally specifying which fields to include or exclude.

Parameters:
  • mode – The mode in which to_python should run. If mode is ‘json’, the output will only contain JSON serializable types. If mode is ‘python’, the output may contain non-JSON-serializable Python objects.

  • include – A set of fields to include in the output.

  • exclude – A set of fields to exclude from the output.

  • context – Additional context to pass to the serializer.

  • by_alias – Whether to use the field’s alias in the dictionary key if defined.

  • exclude_unset – Whether to exclude fields that have not been explicitly set.

  • exclude_defaults – Whether to exclude fields that are set to their default value.

  • exclude_none – Whether to exclude fields that have a value of None.

  • exclude_computed_fields – Whether to exclude computed fields. While this can be useful for round-tripping, it is usually recommended to use the dedicated round_trip parameter instead.

  • round_trip – If True, dumped values should be valid as input for non-idempotent types such as Json[T].

  • warnings – How to handle serialization errors. False/”none” ignores them, True/”warn” logs errors, “error” raises a [PydanticSerializationError][pydantic_core.PydanticSerializationError].

  • fallback – A function to call when an unknown value is encountered. If not provided, a [PydanticSerializationError][pydantic_core.PydanticSerializationError] error is raised.

  • serialize_as_any – Whether to serialize fields with duck-typing serialization behavior.

  • polymorphic_serialization – Whether to use model and dataclass polymorphic serialization for this call.

Returns:

A dictionary representation of the model.

model_dump_json(
*,
indent: int | None = None,
ensure_ascii: bool = False,
include: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
exclude: set[int] | set[str] | Mapping[int, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | Mapping[str, set[int] | set[str] | Mapping[int, IncEx | bool] | Mapping[str, IncEx | bool] | bool] | None = None,
context: Any | None = None,
by_alias: bool | None = None,
exclude_unset: bool = False,
exclude_defaults: bool = False,
exclude_none: bool = False,
exclude_computed_fields: bool = False,
round_trip: bool = False,
warnings: bool | Literal['none', 'warn', 'error'] = True,
fallback: Callable[[Any], Any] | None = None,
serialize_as_any: bool = False,
polymorphic_serialization: bool | None = None,
) str#
!!! abstract “Usage Documentation”

[model_dump_json](../concepts/serialization.md#json-mode)

Generates a JSON representation of the model using Pydantic’s to_json method.

Parameters:
  • indent – Indentation to use in the JSON output. If None is passed, the output will be compact.

  • ensure_ascii – If True, the output is guaranteed to have all incoming non-ASCII characters escaped. If False (the default), these characters will be output as-is.

  • include – Field(s) to include in the JSON output.

  • exclude – Field(s) to exclude from the JSON output.

  • context – Additional context to pass to the serializer.

  • by_alias – Whether to serialize using field aliases.

  • exclude_unset – Whether to exclude fields that have not been explicitly set.

  • exclude_defaults – Whether to exclude fields that are set to their default value.

  • exclude_none – Whether to exclude fields that have a value of None.

  • exclude_computed_fields – Whether to exclude computed fields. While this can be useful for round-tripping, it is usually recommended to use the dedicated round_trip parameter instead.

  • round_trip – If True, dumped values should be valid as input for non-idempotent types such as Json[T].

  • warnings – How to handle serialization errors. False/”none” ignores them, True/”warn” logs errors, “error” raises a [PydanticSerializationError][pydantic_core.PydanticSerializationError].

  • fallback – A function to call when an unknown value is encountered. If not provided, a [PydanticSerializationError][pydantic_core.PydanticSerializationError] error is raised.

  • serialize_as_any – Whether to serialize fields with duck-typing serialization behavior.

  • polymorphic_serialization – Whether to use model and dataclass polymorphic serialization for this call.

Returns:

A JSON string representation of the model.

classmethod model_json_schema(
by_alias: bool = True,
ref_template: str = '#/$defs/{model}',
schema_generator: type[~pydantic.json_schema.GenerateJsonSchema] = <class 'pydantic.json_schema.GenerateJsonSchema'>,
mode: ~typing.Literal['validation',
'serialization'] = 'validation',
*,
union_format: ~typing.Literal['any_of',
'primitive_type_array'] = 'any_of',
) dict[str, Any]#

Generates a JSON schema for a model class.

Parameters:
  • by_alias – Whether to use attribute aliases or not.

  • ref_template – The reference template.

  • union_format

    The format to use when combining schemas from unions together. Can be one of:

    keyword to combine schemas (the default). - ‘primitive_type_array’: Use the [type](https://json-schema.org/understanding-json-schema/reference/type) keyword as an array of strings, containing each type of the combination. If any of the schemas is not a primitive type (string, boolean, null, integer or number) or contains constraints/metadata, falls back to any_of.

  • schema_generator – To override the logic used to generate the JSON schema, as a subclass of GenerateJsonSchema with your desired modifications

  • mode – The mode in which to generate the schema.

Returns:

The JSON schema for the given model class.

classmethod model_parametrized_name(
params: tuple[type[Any], ...],
) str#

Compute the class name for parametrizations of generic classes.

This method can be overridden to achieve a custom naming scheme for generic BaseModels.

Parameters:

params – Tuple of types of the class. Given a generic class Model with 2 type variables and a concrete model Model[str, int], the value (str, int) would be passed to params.

Returns:

String representing the new class where params are passed to cls as type variables.

Raises:

TypeError – Raised when trying to generate concrete names for non-generic models.

model_post_init(context: Any, /) None#

Override this method to perform additional initialization after __init__ and model_construct. This is useful if you want to do some validation that requires the entire model to be initialized.

classmethod model_rebuild(
*,
force: bool = False,
raise_errors: bool = True,
_parent_namespace_depth: int = 2,
_types_namespace: MappingNamespace | None = None,
) bool | None#

Try to rebuild the pydantic-core schema for the model.

This may be necessary when one of the annotations is a ForwardRef which could not be resolved during the initial attempt to build the schema, and automatic rebuilding fails.

Parameters:
  • force – Whether to force the rebuilding of the model schema, defaults to False.

  • raise_errors – Whether to raise errors, defaults to True.

  • _parent_namespace_depth – The depth level of the parent namespace, defaults to 2.

  • _types_namespace – The types namespace, defaults to None.

Returns:

Returns None if the schema is already “complete” and rebuilding was not required. If rebuilding _was_ required, returns True if rebuilding was successful, otherwise False.

classmethod model_validate(
obj: Any,
*,
strict: bool | None = None,
extra: Literal['allow', 'ignore', 'forbid'] | None = None,
from_attributes: bool | None = None,
context: Any | None = None,
by_alias: bool | None = None,
by_name: bool | None = None,
) Self#

Validate a pydantic model instance.

Parameters:
  • obj – The object to validate.

  • strict – Whether to enforce types strictly.

  • extra – Whether to ignore, allow, or forbid extra data during model validation. See the [extra configuration value][pydantic.ConfigDict.extra] for details.

  • from_attributes – Whether to extract data from object attributes.

  • context – Additional context to pass to the validator.

  • by_alias – Whether to use the field’s alias when validating against the provided input data.

  • by_name – Whether to use the field’s name when validating against the provided input data.

Raises:

ValidationError – If the object could not be validated.

Returns:

The validated model instance.

classmethod model_validate_json(
json_data: str | bytes | bytearray,
*,
strict: bool | None = None,
extra: Literal['allow', 'ignore', 'forbid'] | None = None,
context: Any | None = None,
by_alias: bool | None = None,
by_name: bool | None = None,
) Self#
!!! abstract “Usage Documentation”

[JSON Parsing](../concepts/json.md#json-parsing)

Validate the given JSON data against the Pydantic model.

Parameters:
  • json_data – The JSON data to validate.

  • strict – Whether to enforce types strictly.

  • extra – Whether to ignore, allow, or forbid extra data during model validation. See the [extra configuration value][pydantic.ConfigDict.extra] for details.

  • context – Extra variables to pass to the validator.

  • by_alias – Whether to use the field’s alias when validating against the provided input data.

  • by_name – Whether to use the field’s name when validating against the provided input data.

Returns:

The validated Pydantic model.

Raises:

ValidationError – If json_data is not a JSON string or the object could not be validated.

classmethod model_validate_strings(
obj: Any,
*,
strict: bool | None = None,
extra: Literal['allow', 'ignore', 'forbid'] | None = None,
context: Any | None = None,
by_alias: bool | None = None,
by_name: bool | None = None,
) Self#

Validate the given object with string data against the Pydantic model.

Parameters:
  • obj – The object containing string data to validate.

  • strict – Whether to enforce types strictly.

  • extra – Whether to ignore, allow, or forbid extra data during model validation. See the [extra configuration value][pydantic.ConfigDict.extra] for details.

  • context – Extra variables to pass to the validator.

  • by_alias – Whether to use the field’s alias when validating against the provided input data.

  • by_name – Whether to use the field’s name when validating against the provided input data.

Returns:

The validated Pydantic model.

classmethod parse_file(
path: str | Path,
*,
content_type: str | None = None,
encoding: str = 'utf8',
proto: DeprecatedParseProtocol | None = None,
allow_pickle: bool = False,
) Self#
classmethod parse_obj(obj: Any) Self#
classmethod parse_raw(
b: str | bytes,
*,
content_type: str | None = None,
encoding: str = 'utf8',
proto: DeprecatedParseProtocol | None = None,
allow_pickle: bool = False,
) Self#
classmethod schema(
by_alias: bool = True,
ref_template: str = '#/$defs/{model}',
) Dict[str, Any]#
classmethod schema_json(
*,
by_alias: bool = True,
ref_template: str = '#/$defs/{model}',
**dumps_kwargs: Any,
) str#
classmethod update_forward_refs(**localns: Any) None#
classmethod validate(value: Any) Self#
validator validate_attention_dp_config  »  all fields[source]#
model_computed_fields = {}#
model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}#

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

property model_extra: dict[str, Any] | None#

Get extra fields set during validation.

Returns:

A dictionary of extra fields, or None if config.extra is not set to “allow”.

model_fields = {'batching_wait_iters': FieldInfo(annotation=int, required=False, default=10, description='The number of iterations to wait for batching.'), 'enable_balance': FieldInfo(annotation=bool, required=False, default=False, description='Whether to enable balance.'), 'enable_kv_cache_aware_routing': FieldInfo(annotation=bool, required=False, default=False, description="Enable internal KV cache-aware routing for attention DP. When enabled, distributes requests among ranks within a single instance's attention DP group, routing them to the rank with the matching prefix KV cache to reduce redundant prefill computation."), 'kv_cache_routing_account_for_in_transfer': FieldInfo(annotation=bool, required=False, default=False, description="In-transfer load accounting in KV cache-aware routing. When True, requests still streaming KV to the GEN worker (tracked by the PyExecutor AsyncTransferManager but no longer in active_requests) are folded back into the per-rank load reported via RankState. This can improve balance under heavy disagg traffic but inflates num_active_requests reported upstream, which lets the inference loop's idle-fetch wait expire even when no requests are runnable and causes empty fetch cycles. Default False preserves the prior behaviour (fetch blocks when truly idle). Only used when enable_kv_cache_aware_routing is True."), 'kv_cache_routing_cold_start_warmup': FieldInfo(annotation=bool, required=False, default=False, description='Cold-start mitigation in KV cache-aware routing. When True, the first tp_size relaxed requests after router init are round-robined across ranks (bypassing cache-affinity scoring) so every rank caches the shared system prompt before scoring would otherwise pin all traffic to the first warm rank. Only useful when requests share a long system prefix; for diverse-prompt workloads this can scatter requests that would otherwise consolidate on a single warm rank, wasting prefill. Default False preserves pre-warmup routing. Only used when enable_kv_cache_aware_routing is True.'), 'kv_cache_routing_conversation_affinity': FieldInfo(annotation=bool, required=False, default=False, description="Enable explicit conversation-affinity routing for attention DP. When True, the first request of each conversation is round-robined across ranks and every subsequent request carrying the same conversation_params.conversation_id is pinned to that conversation's first-turn rank. OpenAI requests use the body conversation_params as canonical; the serve edge only creates conversation_params from the X-Session-ID header when the body does not provide it. This keeps a multi-turn conversation's KV-cache prefix on one rank (maximizing block reuse, minimizing cross-rank migration). Unlike enable_kv_cache_aware_routing (affinity inferred from prefix-match length, which is lost when blocks are evicted), the conversation->rank map is explicit and survives eviction. Falls back to load-balanced round-robin when no conversation_id is available. Takes precedence over enable_kv_cache_aware_routing when both are set."), 'kv_cache_routing_fair_share_multiplier': FieldInfo(annotation=float, required=False, default=2.0, description='Loose per-rank active-request cap in KV cache-aware routing, expressed as a multiplier of the ceil fair-share (ceil((total_active + new) / tp_size)). Once a rank hits this cap within a scheduling batch it is removed from the eligible set for the remainder of the batch. Default 2.0 permits a 2x slack so cache affinity can dominate while preventing runaway concentration on a single rank. Set to 1.0 for strict fair share. Only used when enable_kv_cache_aware_routing is True.'), 'kv_cache_routing_load_balance_weight': FieldInfo(annotation=float, required=False, default=1.0, description='Weight (beta) for the load-balance term in KV cache-aware routing. Higher values prioritize load balance over cache affinity. Only used when enable_kv_cache_aware_routing is True.'), 'kv_cache_routing_match_rate_threshold': FieldInfo(annotation=float, required=False, default=0.1, description='Cache-affinity gate in KV cache-aware routing. For each request, match_len contributes to scoring only when max(match_len) / request_tokens across eligible ranks is strictly above this threshold; otherwise match_len is forced to 0 so routing is driven purely by load. Default 0.1 requires at least a 10% hit rate before cache affinity kicks in, which prevents a small universal prefix (e.g. a shared system prompt) from pinning all traffic to the first warm ranks. Set to 0.0 to honour any nonzero match. Only used when enable_kv_cache_aware_routing is True.'), 'kv_cache_routing_max_sessions': FieldInfo(annotation=int, required=False, default=65536, description='LRU cap on the conversation->rank map used by conversation-affinity routing. The oldest conversations are evicted once more than this many are tracked, bounding memory on long-running servers. Only used when kv_cache_routing_conversation_affinity is True.'), 'kv_cache_routing_new_conv_placement': FieldInfo(annotation=Literal['round_robin', 'least_queued'], required=False, default='round_robin', description="Placement policy in conversation-affinity routing for requests with no pinned rank yet (first turn of a conversation, requests without a conversation_id, sticky overflow). 'round_robin' (default) equalizes per-rank conversation counts. 'least_queued' places them on the rank with the fewest live requests instead: per-conversation load (turn rate, fan-out, prefill length) is not uniform, so count-uniform round-robin can leave some ranks with deep queues while others idle; steering new conversations by queue depth evens that out and cuts tail TTFT. Existing conversation->rank pins are unaffected. Only used when kv_cache_routing_conversation_affinity is True."), 'timeout_iters': FieldInfo(annotation=int, required=False, default=50, description='The number of iterations to timeout.')}#
property model_fields_set: set[str]#

Returns the set of fields that have been explicitly set on this model instance.

Returns:

A set of strings representing the fields that have been set,

i.e. that were not filled from defaults.