AttentionDpConfig#

class tensorrt_llm.llmapi.AttentionDpConfig(
*,
enable_balance: bool = False,
timeout_iters: int = 50,
batching_wait_iters: int = 10,
enable_kv_cache_aware_routing: bool = False,
kv_cache_routing_load_balance_weight: float = 1.0,
kv_cache_routing_match_rate_threshold: float = 0.1,
kv_cache_routing_fair_share_multiplier: float = 2.0,
kv_cache_routing_cold_start_warmup: bool = False,
kv_cache_routing_account_for_in_transfer: bool = False,
kv_cache_routing_conversation_affinity: bool = False,
kv_cache_routing_max_sessions: int = 65536,
kv_cache_routing_new_conv_placement: Literal['round_robin', 'least_queued', 'least_tokens'] = 'round_robin',
)[source]#

Bases: StrictBaseModel

Configuration for attention DP.

field batching_wait_iters: int = 10#

The number of iterations to wait for batching.

field enable_balance: bool = False#

Whether to enable balance.

field enable_kv_cache_aware_routing: bool = False#

Enable internal KV cache-aware routing for attention DP. When enabled, distributes requests among ranks within a single instance’s attention DP group, routing them to the rank with the matching prefix KV cache to reduce redundant prefill computation.

field kv_cache_routing_account_for_in_transfer: bool = False#

In-transfer load accounting in KV cache-aware routing. When True, requests still streaming KV to the GEN worker (tracked by the PyExecutor AsyncTransferManager but no longer in active_requests) are folded back into the per-rank load reported via RankState. This can improve balance under heavy disagg traffic but inflates num_active_requests reported upstream, which lets the inference loop’s idle-fetch wait expire even when no requests are runnable and causes empty fetch cycles. Default False preserves the prior behaviour (fetch blocks when truly idle). Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_cold_start_warmup: bool = False#

Cold-start mitigation in KV cache-aware routing. When True, the first tp_size relaxed requests after router init are round-robined across ranks (bypassing cache-affinity scoring) so every rank caches the shared system prompt before scoring would otherwise pin all traffic to the first warm rank. Only useful when requests share a long system prefix; for diverse-prompt workloads this can scatter requests that would otherwise consolidate on a single warm rank, wasting prefill. Default False preserves pre-warmup routing. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_conversation_affinity: bool = False#

Enable explicit conversation-affinity routing for attention DP. When True, the first request is placed by kv_cache_routing_new_conv_placement and every subsequent request carrying the same conversation_params.conversation_id is pinned to that conversation’s first-turn rank. OpenAI requests use the body conversation_params as canonical; the serve edge only creates conversation_params from the X-Session-ID header when the body does not provide it. This keeps a multi-turn conversation’s KV-cache prefix on one rank (maximizing block reuse, minimizing cross-rank migration). Unlike enable_kv_cache_aware_routing (affinity inferred from prefix-match length, which is lost when blocks are evicted), the conversation->rank map is explicit and survives eviction. The same placement policy is used when no conversation_id is available. Takes precedence over enable_kv_cache_aware_routing when both are set.

field kv_cache_routing_fair_share_multiplier: float = 2.0#

Loose per-rank active-request cap in KV cache-aware routing, expressed as a multiplier of the ceil fair-share (ceil((total_active + new) / tp_size)). Once a rank hits this cap within a scheduling batch it is removed from the eligible set for the remainder of the batch. Default 2.0 permits a 2x slack so cache affinity can dominate while preventing runaway concentration on a single rank. Set to 1.0 for strict fair share. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_load_balance_weight: float = 1.0#

Weight (beta) for the load-balance term in KV cache-aware routing. Higher values prioritize load balance over cache affinity. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_match_rate_threshold: float = 0.1#

Cache-affinity gate in KV cache-aware routing. For each request, match_len contributes to scoring only when max(match_len) / request_tokens across eligible ranks is strictly above this threshold; otherwise match_len is forced to 0 so routing is driven purely by load. Default 0.1 requires at least a 10% hit rate before cache affinity kicks in, which prevents a small universal prefix (e.g. a shared system prompt) from pinning all traffic to the first warm ranks. Set to 0.0 to honour any nonzero match. Only used when enable_kv_cache_aware_routing is True.

field kv_cache_routing_max_sessions: int = 65536#

LRU cap on the conversation->rank map used by conversation-affinity routing. The oldest conversations are evicted once more than this many are tracked, bounding memory on long-running servers. Only used when kv_cache_routing_conversation_affinity is True.

field kv_cache_routing_new_conv_placement: Literal['round_robin', 'least_queued', 'least_tokens'] = 'round_robin'#

Placement policy in conversation-affinity routing for requests with no pinned rank yet (first turn of a conversation, requests without a conversation_id, sticky overflow). ‘round_robin’ (default) rotates across eligible ranks. ‘least_queued’ chooses the fewest live requests. ‘least_tokens’ chooses the fewest active prompt tokens, including requests assigned in the current batch; ties use live request count, then the rotating rank cursor. Active prompt tokens estimate work, not cached KV residency. Existing conversation->rank pins are unaffected. Only used when kv_cache_routing_conversation_affinity is True.

field timeout_iters: int = 50#

The number of iterations to timeout.

__init__(**data: Any) → None#

Create a new model by parsing and validating input data from keyword arguments.

Raises [ValidationError][pydantic_core.ValidationError] if the input data cannot be validated to form a valid model.

self is explicitly positional-only to allow self as a field name.

validator validate_attention_dp_config  »  all fields[source]#