Runtime Stepper#
-
class RuntimeStepper#
The runtime’s execution facade for the stepped control plane: typed operations over immutable inputs, typed results, no callbacks, and no mutable state lent across the boundary.
Wraps a live GenerationSession. In this stage the facade is exercised by unit tests while the engine still drives the loop through handleRequest; the loop moves behind this interface in the next stage, at which point GenerationBoundaryHook is deleted.
Threading: one thread drives a stepper, the same thread that owns the session.
Public Functions
-
explicit RuntimeStepper(LLMRankRuntime::GenerationSession &session)#
Adopts the session’s canonical resident identities.
-
RuntimeStepper(RuntimeStepper const&) = delete#
-
RuntimeStepper &operator=(RuntimeStepper const&) = delete#
-
AdmissionResult admit(AdmissionIntent intent)#
Logical admission only: reservation (budget + context-cache lease) and the seat commit at appendSlot. The seated prefill is not run here — it is the next prefill tick, and must be the next operation on this stepper (the admission staging buffers stay valid exactly until the next admission’s preprocess).
-
StepResult prefill(ImmutablePrefillBatch const &batch)#
The seated batch-1 prefill for the pending admission (batch.target must be its ref), or — when nothing is pending — the founding prefill of the wrapped session.
-
StepResult decode(ImmutableDecodeBatch const &batch)#
One decode step over the live batch. batch.residents must name every live resident: this stage runs a single cohort, and the view exists so the executed set is an explicit input.
-
std::vector<ResidentRef> residents() const#
The refs of every live resident, in slot order. The caller’s view for building batches.
-
explicit RuntimeStepper(LLMRankRuntime::GenerationSession &session)#
-
class SteppedExecution#
What the scheduler drives, one typed operation per tick, however many ranks execute it. The single-rank implementation wraps one stepper directly; the thread-parallel implementation broadcasts each operation as a command every rank executes collectively. Either way the scheduler sees the same vocabulary and the runtime never calls back into policy.
Subclassed by trt_edgellm::rt::SteppedRequest
Public Functions
-
virtual ~SteppedExecution() = default#
-
virtual std::vector<ResidentRef> residents() const = 0#
The refs of every live resident, in slot order (identical on every rank by determinism).
- virtual AdmissionResult admit(
- LLMGenerationRequest const &request,
- int32_t originalIndex,
- RequestId requestId
Logical admission of
requestatoriginalIndex:intent building (validation, tokenization, encoder preprocess) plus the seat commit. The request itself is the cross-rank payload — every rank rebuilds the intent deterministically, exactly the property the boundary relay relied on. The seated prefill is the next prefill tick.
-
virtual StepResult prefill(ImmutablePrefillBatch const &batch) = 0#
-
virtual StepResult decode(ImmutableDecodeBatch const &batch) = 0#
- virtual LLMGenerationResponse materialize(
- BatchResult const &result,
- std::vector<std::string> const &stopStrings
Convert a finished snapshot to a response (rank zero’s tokenizer and stop trim).
-
virtual bool finish(LLMGenerationResponse &response) = 0#
Everything handleRequest did after the loop. Call once, when no residents remain.
-
virtual void abort() noexcept = 0#
Abandon an incomplete session and release all runtime-owned request state.
-
virtual ~SteppedExecution() = default#
-
class SteppedRequest : public trt_edgellm::rt::SteppedExecution#
One request held open under the stepped control plane: the prepared request, the runtime-side generation state, the session, and the stepper over it — everything handleRequest kept on its stack, owned across ticks instead. Created by RuntimeCoordinator::beginStepped(); the single-rank SteppedExecution.
Public Functions
-
inline RuntimeStepper &stepper() noexcept#
-
virtual std::vector<ResidentRef> residents() const override#
The refs of every live resident, in slot order (identical on every rank by determinism).
- virtual AdmissionResult admit(
- LLMGenerationRequest const &request,
- int32_t originalIndex,
- RequestId requestId
Logical admission of
requestatoriginalIndex:intent building (validation, tokenization, encoder preprocess) plus the seat commit. The request itself is the cross-rank payload — every rank rebuilds the intent deterministically, exactly the property the boundary relay relied on. The seated prefill is the next prefill tick.
- virtual StepResult prefill(
- ImmutablePrefillBatch const &batch
-
virtual StepResult decode(ImmutableDecodeBatch const &batch) override#
- virtual LLMGenerationResponse materialize(
- BatchResult const &result,
- std::vector<std::string> const &stopStrings
Convert a finished snapshot to a response (rank zero’s tokenizer and stop trim).
-
virtual bool finish(LLMGenerationResponse &response) override#
Everything handleRequest did after the loop. Call once, when no residents remain.
-
virtual void abort() noexcept override#
Abandon an incomplete session and release all runtime-owned request state.
- AdmissionIntent buildIntent(
- LLMGenerationRequest const &request,
- int32_t originalIndex,
- RequestId requestId
Stage-one intent building for a joiner (validation, tokenization, encoder preprocess), against this request’s live session. Throws for requests the batch rejects outright.
-
~SteppedRequest() override#
Public Static Functions
- static std::unique_ptr<SteppedRequest> begin(
- LLMRankRuntime &runtime,
- LLMGenerationRequest prepared,
- RequestId requestId,
- cudaStream_t stream,
- LLMRankRuntime::TokenBroadcastFn tokenBroadcast = nullptr,
- int32_t parallelRank = -1
Runs everything up to and including the founding prefill (LLMRankRuntime::beginGeneration) and wires the session and stepper over the result. Returns null on the same refusals that made handleRequest return false; throws where it threw.
preparedis the coordinator’s prepared request state and is owned here because the session references it across ticks.
-
inline RuntimeStepper &stepper() noexcept#
-
struct AdmissionIntent#
Everything validation, tokenization, and the encoder preprocess produce. Building one touches no resident state; discarding one leaves the batch byte-identical.
-
struct AdmissionResult#
Public Types
-
struct ImmutablePrefillBatch#
The immutable input to one executed step: which residents participate, materialized by the caller. The prefill-first scheduler runs a single cohort, so a decode view is “every live resident” and a prefill view is the one newly admitted ref; the types exist so the executed set is an explicit input rather than stepper-internal knowledge, and so B=1/B=N views can evolve without an API change.
Public Members
-
ResidentRef target#
-
ResidentRef target#
-
struct ImmutableDecodeBatch#
Public Members
-
std::vector<ResidentRef> residents#
-
std::vector<ResidentRef> residents#
-
struct TokenDelta#
Tokens one resident produced in one step.
Public Members
-
std::vector<int32_t> tokenIds#
-
std::vector<int32_t> tokenIds#
-
struct StepResult#
What one executed step did. The single commit unit: everything in it happened together, and nothing outside it happened.
Public Members
-
bool ok = {false}#
-
SequenceWork work = {SequenceWork::kContext}#
-
StepId stepId = {0}#
Runtime-assigned identity of the ScheduledStep whose effects this result commits.
-
std::vector<SequenceIdentity> participants#
Exact ordered logical/physical cohort executed by this step.
-
std::vector<ScheduledSequenceDescriptor> execution#
Pointer-free logical execution plan used for cross-rank shape/state consensus.
-
std::vector<std::pair<ResidentRef, TokenDelta>> deltas#
Tokens produced this step, per participating resident (absent for a resident that produced none).
-
std::vector<std::pair<ResidentRef, BatchResult>> finished#
Residents that reached a terminal state and were evicted this step, with the finished sequence snapshot; their physical resident slots are released (and their refs invalidated) as part of this commit unit. A seat that failed its seated prefill surfaces here on the step whose eviction files it, not before. Dense-row compaction is runtime-private and never reported: refs are stable resident identities, so the caller’s table needs no re-keying.
-
bool publishedPrefix = {false}#
True when this step published context-cache prefix blocks (cache visibility commits with the step that produced it, never ahead of it).
-
bool ok = {false}#