Sub-agent Routing#
Sub-agent routing lets a sub-agent reuse its parent’s placement when conversation-aware affinity routing is enabled. Co-locating their requests on the same context instance and attention data-parallel (ADP) rank improves the opportunity to reuse shared prompt prefixes. Each agent retains its own conversation ID for KV-cache and conversation management, which track a linear history per conversation.
Configuration#
Add the following settings to an existing disaggregated serving configuration. Merge the router settings into the corresponding server sections:
conversation_affinity_header_for_subagents: X-Dynamo-Parent-Session-ID
subagent_affinity_scope: context
internal_request_auth_key: <shared-secret>
context_servers:
router:
type: conversation
The feature is opt-in: leaving conversation_affinity_header_for_subagents unset
preserves ordinary conversation routing. Choose a dedicated parent-session header
supplied by your agent gateway, separate from the headers identifying each agent’s
own conversation.
Set the same internal_request_auth_key on the edge and its context/generation
workers. A combined disaggregated launch propagates this key to its workers;
independently launched workers need it in their own configuration. Forwarding
sub-agent affinity requires this key, even while KV-transfer authentication still
permits a missing key with a transitional warning. The edge rejects an affinity
configuration without an internal key at startup.
Workers may omit --server_role: when no role is configured, affinity
validation infers context or generation from disaggregated_params.request_type.
The inferred role selects the signature to verify; a valid signature is still
required. An explicitly configured role takes precedence.
On the context workers, enable conversation affinity in your existing attention-DP configuration:
enable_attention_dp: true
attention_dp_config:
kv_cache_routing_conversation_affinity: true
For token-aware placement, add
kv_cache_routing_new_conv_placement: least_tokens under attention_dp_config.
This chooses an eligible ADP rank with the lowest active prompt-token load,
including input tokens of requests assigned in the current batch. The default
is round_robin; least_queued uses active-request counts. Existing
conversation affinity still applies.
Use /v1/chat/completions for both instance and ADP-rank affinity. The
/v1/completions endpoint supports instance affinity only.
Request headers#
Send a stable, distinct x-session-id for each agent. On sub-agent requests, also
send the configured parent header with the parent’s conversation ID. Main-agent
requests carry their own session ID only. An explicit body
conversation_params.conversation_id takes precedence over session headers.
For example, with X-Dynamo-Parent-Session-ID configured:
Request |
|
|
Context routing key |
Conversation ID |
|---|---|---|---|---|
Parent |
|
Omitted |
|
|
Sub-agent A |
|
|
|
|
Sub-agent B |
|
|
|
|
Keep these headers stable across each sub-agent’s turns. This feature supports one level of sub-agents; nested sub-agent routing is not supported.
Routing scope#
subagent_affinity_scope controls which fleets use the parent routing key:
Scope |
Context fleet |
Generation fleet |
|---|---|---|
|
Parent affinity |
Child’s own conversation routing |
|
Parent affinity |
Parent affinity |
The default scope targets shared prefill cache reuse while allowing children to
spread across generation workers. To use both, configure the generation fleet
with router.type: conversation and enable conversation-aware ADP affinity on its
workers as well. Use context scope with conditional disaggregation, which requires
a KV-cache-aware generation router.
Placement follows the existing conversation routers’ behavior. Affinity is best effort: a saturated ADP rank can overflow to another rank, and affinity bindings can be evicted. Cache reuse depends on matching prompt prefixes and cache availability.
Gateway trust boundary#
The parent-session header is routing input, not proof of a parent-child relationship. TensorRT-LLM does not authenticate that relationship: a client that can set the configured header can request affinity to another session’s instance and ADP rank.
Deployments accepting untrusted requests should place a gateway before
trtllm-serve. The gateway should prune client-supplied copies of the configured
parent-session header and the internal x-trtllm-subagent-affinity-id and
x-trtllm-subagent-affinity-auth headers,
then set the parent-session header only for authorized sub-agent requests.
Keep worker endpoints behind this trusted boundary as well. Each agent’s own
session ID should continue to identify its independent conversation.
How it works#
The disagg edge reads the configured parent header into an internal routing key.
The conversation router uses that key to select the parent’s instance, and the HTTP
client forwards it as x-trtllm-subagent-affinity-id with a separate HMAC signature
in x-trtllm-subagent-affinity-auth. Context/generation workers validate the
signature before placing the key in SchedulingParams.subagent_affinity_id for
the conversation-aware ADP router. The signature binds the hint to the worker
role, model, child’s conversation ID, request type, and disaggregated request ID.
Ordinary aggregated requests with neither a configured worker role nor a
context/generation request type ignore the affinity header and retain ordinary
conversation routing. Worker requests carrying unsigned or invalid affinity are
rejected, including requests sent to workers without an explicit role.
The child’s conversation ID continues to identify its own history throughout this path. The internal affinity field is excluded from the serialized request body; older workers can ignore the new header during a rolling deployment. Upgrade edges before workers: upgraded workers reject affinity headers from older edges that do not sign them. The existing KV-transfer authentication signature remains unchanged.