Request Record#

struct RequestRecord#

Everything about one in-flight request that both the caller and the actor need to see.

Held by shared_ptr on both sides on purpose: the caller may drop its RequestHandle while the actor is still mid-request, and the actor must still have somewhere to put the outcome. It also means a cancel arriving after retirement finds nothing rather than dangling.

Public Functions

inline void publishOutcome(
TerminalStatus terminalStatus,
std::string message = {}
)#

Actor: publish the outcome and release anyone waiting for it.

inline void wakeWaiters()#

Wake a get() parked on the outcome without publishing one: it re-evaluates the channel’s cancel flag, which it treats as terminal for an unfinished request. The empty critical section orders the wake after the waiter’s predicate check.

Public Members

std::chrono::steady_clock::time_point submittedAt = {}#

When submit() accepted the request; the queue-latency metric measures from here to the moment the actor starts serving it (founding or mid-flight admission).

RequestId id = {kInvalidRequestId}#
std::shared_ptr<StreamChannel> channel#

The channel the runtime itself writes tokens into, handed straight to the caller.

Deliberately the runtime’s own type rather than a scheduler-owned one. The runtime already delivers per slot: it pushes through slotStreams[i].channel, reads cancellation from the same place at the top of every step, and moves the entry when the batch is compacted. A separate channel would have to be bridged onto this one for no gain.

std::shared_ptr<StreamChannel> runtimeChannel#

The channel the runtime actually polls for the cancel flag: the caller’s own first channel when the request brought one, this record’s otherwise. Cancellation must set the flag here or a caller-supplied channel would leave a running request uncancellable.

LLMGenerationResponse response#
TerminalStatus status = {TerminalStatus::kCompleted}#
std::string errorMessage#

Set only when status is kExecutionError, so get() can rethrow something actionable.

std::mutex outcomeMutex#

Publishes the three fields above, and wakes a caller waiting on them.

This, not the channel’s terminal flag, is what says a request is over. The two answer different questions and are set by different parties: the runtime finishes the channel when the token stream ends, while the actor records the outcome after execution returns &#8212; and for a request that never reached the runtime at all, the channel is never touched. Treating a terminal channel as “the outcome is readable” would sometimes read a response that had not been written, and would leave a discarded request waiting forever.

std::condition_variable outcomeReady#
std::atomic<bool> outcomePublished = {false}#