Runtime Stepper#

class RuntimeStepper#

The runtime’s execution facade for the stepped control plane: typed operations over immutable inputs, typed results, no callbacks, and no mutable state lent across the boundary.

Wraps a live GenerationSession. In this stage the facade is exercised by unit tests while the engine still drives the loop through handleRequest; the loop moves behind this interface in the next stage, at which point GenerationBoundaryHook is deleted.

Threading: one thread drives a stepper, the same thread that owns the session.

Public Functions

explicit RuntimeStepper(LLMRankRuntime::GenerationSession &session)#

Adopts the session’s canonical resident identities.

RuntimeStepper(RuntimeStepper const&) = delete#
RuntimeStepper &operator=(RuntimeStepper const&) = delete#
AdmissionResult admit(AdmissionIntent intent)#

Logical admission only: reservation (budget + context-cache lease) and the seat commit at appendSlot. The seated prefill is not run here — it is the next prefill tick, and must be the next operation on this stepper (the admission staging buffers stay valid exactly until the next admission’s preprocess).

StepResult prefill(ImmutablePrefillBatch const &batch)#

The seated batch-1 prefill for the pending admission (batch.target must be its ref), or — when nothing is pending — the founding prefill of the wrapped session.

StepResult decode(ImmutableDecodeBatch const &batch)#

One decode step over the live batch. batch.residents must name every live resident: this stage runs a single cohort, and the view exists so the executed set is an explicit input.

std::vector<ResidentRef> residents() const#

The refs of every live resident, in slot order. The caller’s view for building batches.

class SteppedExecution#

What the scheduler drives, one typed operation per tick, however many ranks execute it. The single-rank implementation wraps one stepper directly; the thread-parallel implementation broadcasts each operation as a command every rank executes collectively. Either way the scheduler sees the same vocabulary and the runtime never calls back into policy.

Subclassed by trt_edgellm::rt::SteppedRequest

Public Functions

virtual ~SteppedExecution() = default#
virtual std::vector<ResidentRef> residents() const = 0#

The refs of every live resident, in slot order (identical on every rank by determinism).

virtual AdmissionResult admit(
LLMGenerationRequest const &request,
int32_t originalIndex,
RequestId requestId
) = 0#

Logical admission of request at originalIndex: intent building (validation, tokenization, encoder preprocess) plus the seat commit. The request itself is the cross-rank payload &#8212; every rank rebuilds the intent deterministically, exactly the property the boundary relay relied on. The seated prefill is the next prefill tick.

virtual StepResult prefill(ImmutablePrefillBatch const &batch) = 0#
virtual StepResult decode(ImmutableDecodeBatch const &batch) = 0#
virtual LLMGenerationResponse materialize(
BatchResult const &result,
std::vector<std::string> const &stopStrings
) const = 0#

Convert a finished snapshot to a response (rank zero’s tokenizer and stop trim).

virtual bool finish(LLMGenerationResponse &response) = 0#

Everything handleRequest did after the loop. Call once, when no residents remain.

virtual void abort() noexcept = 0#

Abandon an incomplete session and release all runtime-owned request state.

class SteppedRequest : public trt_edgellm::rt::SteppedExecution#

One request held open under the stepped control plane: the prepared request, the runtime-side generation state, the session, and the stepper over it &#8212; everything handleRequest kept on its stack, owned across ticks instead. Created by RuntimeCoordinator::beginStepped(); the single-rank SteppedExecution.

Public Functions

inline RuntimeStepper &stepper() noexcept#
virtual std::vector<ResidentRef> residents() const override#

The refs of every live resident, in slot order (identical on every rank by determinism).

virtual AdmissionResult admit(
LLMGenerationRequest const &request,
int32_t originalIndex,
RequestId requestId
) override#

Logical admission of request at originalIndex: intent building (validation, tokenization, encoder preprocess) plus the seat commit. The request itself is the cross-rank payload &#8212; every rank rebuilds the intent deterministically, exactly the property the boundary relay relied on. The seated prefill is the next prefill tick.

virtual StepResult prefill(
ImmutablePrefillBatch const &batch
) override#
virtual StepResult decode(ImmutableDecodeBatch const &batch) override#
virtual LLMGenerationResponse materialize(
BatchResult const &result,
std::vector<std::string> const &stopStrings
) const override#

Convert a finished snapshot to a response (rank zero’s tokenizer and stop trim).

virtual bool finish(LLMGenerationResponse &response) override#

Everything handleRequest did after the loop. Call once, when no residents remain.

virtual void abort() noexcept override#

Abandon an incomplete session and release all runtime-owned request state.

AdmissionIntent buildIntent(
LLMGenerationRequest const &request,
int32_t originalIndex,
RequestId requestId
)#

Stage-one intent building for a joiner (validation, tokenization, encoder preprocess), against this request’s live session. Throws for requests the batch rejects outright.

~SteppedRequest() override#

Public Static Functions

static std::unique_ptr<SteppedRequest> begin(
LLMRankRuntime &runtime,
LLMGenerationRequest prepared,
RequestId requestId,
cudaStream_t stream,
LLMRankRuntime::TokenBroadcastFn tokenBroadcast = nullptr,
int32_t parallelRank = -1
)#

Runs everything up to and including the founding prefill (LLMRankRuntime::beginGeneration) and wires the session and stepper over the result. Returns null on the same refusals that made handleRequest return false; throws where it threw. prepared is the coordinator’s prepared request state and is owned here because the session references it across ticks.

struct AdmissionIntent#

Everything validation, tokenization, and the encoder preprocess produce. Building one touches no resident state; discarding one leaves the batch byte-identical.

Public Members

uint64_t requestId = {0}#

Engine-level identity, carried through for the caller’s bookkeeping; the runtime keys residents by the seed’s originalIndex.

SlotSeed seed#
struct AdmissionResult#

Public Types

enum class Status : uint8_t#

Values:

enumerator kAdmitted#
enumerator kNoCapacity#

Transient pressure; nothing was modified, retry at a later tick.

enumerator kRejected#

The batch rejects the intent outright; nothing was modified.

Public Members

Status status = {Status::kRejected}#
ResidentRef ref#

valid iff kAdmitted

std::string reason#

diagnostic for kRejected

struct ImmutablePrefillBatch#

The immutable input to one executed step: which residents participate, materialized by the caller. The prefill-first scheduler runs a single cohort, so a decode view is “every live resident” and a prefill view is the one newly admitted ref; the types exist so the executed set is an explicit input rather than stepper-internal knowledge, and so B=1/B=N views can evolve without an API change.

Public Members

ResidentRef target#
struct ImmutableDecodeBatch#

Public Members

std::vector<ResidentRef> residents#
struct TokenDelta#

Tokens one resident produced in one step.

Public Members

std::vector<int32_t> tokenIds#
struct StepResult#

What one executed step did. The single commit unit: everything in it happened together, and nothing outside it happened.

Public Members

bool ok = {false}#
SequenceWork work = {SequenceWork::kContext}#
StepId stepId = {0}#

Runtime-assigned identity of the ScheduledStep whose effects this result commits.

std::vector<SequenceIdentity> participants#

Exact ordered logical/physical cohort executed by this step.

std::vector<ScheduledSequenceDescriptor> execution#

Pointer-free logical execution plan used for cross-rank shape/state consensus.

std::vector<std::pair<ResidentRef, TokenDelta>> deltas#

Tokens produced this step, per participating resident (absent for a resident that produced none).

std::vector<std::pair<ResidentRef, BatchResult>> finished#

Residents that reached a terminal state and were evicted this step, with the finished sequence snapshot; their physical resident slots are released (and their refs invalidated) as part of this commit unit. A seat that failed its seated prefill surfaces here on the step whose eviction files it, not before. Dense-row compaction is runtime-private and never reported: refs are stable resident identities, so the caller’s table needs no re-keying.

bool publishedPrefix = {false}#

True when this step published context-cache prefix blocks (cache visibility commits with the step that produced it, never ahead of it).