Batch Compatibility#
-
class BatchCompatibility#
Whether two requests may occupy the same forward pass.
Some request fields configure the engine for the whole step rather than for one sequence: a single temperature is applied to the batch, one set of LoRA weights is bound, one decoder is selected. Two requests that disagree on any of those cannot be batched, and admitting them together would silently give one of them the other’s settings.
The classification is closed by convention: every field of LLMGenerationRequest is either in the compared list or in the exempt list beside it, and a new field must be placed in one of the two when it is added — LLMGenerationRequest’s own declaration points here for that reason. The failure mode this guards is silent staleness: a new batch-relevant field that nobody classifies is a check that keeps passing while one request quietly generates with another’s settings.
Exempt fields fall into two groups, and nothing else is exempt:
per-sequence payload: the prompts themselves, their pre-tokenized form, and the stream each result is delivered on;
preprocessing that has already happened by the time a request is admitted, so it cannot affect a shared step: chat templating and generation-prompt insertion. Audio generation and hidden-state capture are NOT exempt: they are compared (see the list in batchCompatibility.cpp), so enabling one never silently shares a step with a request that has not.
Public Static Functions
- static bool compatible(
- LLMGenerationRequest const &resident,
- LLMGenerationRequest const &candidate
True if
candidatecan join a step that is already runningresident.
- static std::string firstDifference(
- LLMGenerationRequest const &resident,
- LLMGenerationRequest const &candidate
The first field that differs, for diagnostics. Empty when the two are compatible.
Reported by name because “incompatible request” alone leaves the caller guessing which of twenty fields it got wrong.