Context Cache Request#

class ContextCacheRequest#

Owns one admitted runtime request and binds it to the context-cache coordinator lifecycle.

Public Types

enum class AdmitSequenceStatus : uint8_t#

Values:

enumerator kAdmitted#
enumerator kNoCapacity#

Transient pool pressure: pages free as resident sequences retire, so the caller may retry at a later step boundary rather than failing the request.

enumerator kFailed#

Public Functions

ContextCacheRequest(ContextCacheRequest&&) noexcept = default#
ContextCacheRequest &operator=(ContextCacheRequest&&) = delete#
ContextCacheRequest(ContextCacheRequest const&) = delete#
ContextCacheRequest &operator=(ContextCacheRequest const&) = delete#
~ContextCacheRequest() = default#
std::vector<int32_t> const &prefillStarts() const noexcept#

Per-sequence logical offsets at which runtime prefill begins.

int32_t reuseTokenLength(int32_t slot) const noexcept#

Number of reused prefix tokens for a slot (== the logical prefill start offset).

AdmitSequenceStatus admitSequence(
std::vector<int32_t> const &tokenIds,
std::string const &loraWeightsName,
DecodingKvHeadroom const &headroom,
int32_t &prefillStart,
ResidentRef resident,
cudaStream_t stream,
std::vector<int32_t> const &mediaTokenIds = {},
std::vector<imageUtils::ImageData> const &imageBuffers = {},
std::vector<audioUtils::AudioData> const &audioBuffers = {}
)#

Join one more text-only sequence to this live request: lookup, lease, and row binding. On kAdmitted, prefillStart receives the reused prefix length the seated prefill skips.

bool finalizeSequenceAdmission(
int32_t slot,
int32_t const &lookaheadToken,
int32_t fullInputLength
)#

Record the seated prefill’s lookahead token so the slot’s ledger matches its pages, and publish the ready prefix blocks. Call once, right after the seated prefill succeeds.

bool retractSequenceAdmission() noexcept#

Undo the most recent admitSequence before its slot ever joined the runtime batch: the recovery path for a seating that threw between lease and slot append.

bool publishHybridMtpEndpoint(
int32_t slot,
int32_t residentStateLength,
Tensor const &baseHiddenStates,
int32_t boundaryHiddenRow
)#

Publish one Hybrid+MTP checkpoint at the stable predecessor boundary. Forwards to the coordinator’s dedicated MTP publication entrypoint; the runtime drives this after the folded draft prefill materialized boundary state.

bool restoreHybridMtpBoundaryHidden(
int32_t slot,
Tensor &baseHiddenStates,
int32_t destinationRow
)#

Restore the reused checkpoint’s saved boundary base-hidden row into baseHiddenStates for the fold micro-forward.

bool preparePrefill()#
bool enqueuePrefillCaptures()#
bool completePrefill(
DecodingInferenceContext const &context,
std::vector<int32_t> const &commonStateLengths
)#
bool prepareDecodeStep(
DecodingInferenceContext const &context,
DecodingKvHeadroom const &headroom
)#
bool completeDecodeStep(
DecodingInferenceContext const &context,
std::vector<int32_t> const &commonStateLengths
)#
bool beginBatchCompaction(
std::vector<int32_t> const &oldToNew,
int32_t newBatchSize,
Tensor &deviceBatchMapping
)#
bool completeBatchCompaction(
std::vector<int32_t> const &keepMapping
)#

Compact the per-slot reuse bookkeeping to survivors. keepMapping is the runtime’s oldSlot -> newSlot batch mapping (-1 for an evicted slot), the same one performBatchEvict uses; without it the reuse-length vector desyncs from the batch after the first eviction.

bool finish()#

Public Static Functions

static std::optional<ContextCacheRequest> begin(
ContextCacheCoordinator &coordinator,
LLMGenerationRequest const &request,
DecodingInferenceContext const &context,
bool speculativeRequest,
DecodingKvHeadroom const &headroom,
std::vector<int32_t> const &mediaTokenIds = {},
DecodingTokenStateContract tokenStateContract = DecodingTokenStateContract::kCommittedPlusLookahead,
ContextCacheCommitPolicy commitPolicy = ContextCacheCommitPolicy::kIncludingGeneratedTokens
)#

Admit one tokenized request and bind its cache resources. A disengaged result means admission failed.

Parameters:

mediaTokenIds – Placeholder token IDs for media modalities (e.g. image, audio). Positions matching any of these IDs are content-hashed for cache differentiation.

Throws:

std::runtime_error – if a media position has to be hashed and its pixels are not readable on the host.