Context Cache Request#
-
class ContextCacheRequest#
Owns one admitted runtime request and binds it to the context-cache coordinator lifecycle.
Public Types
Public Functions
-
ContextCacheRequest(ContextCacheRequest&&) noexcept = default#
-
ContextCacheRequest &operator=(ContextCacheRequest&&) = delete#
-
ContextCacheRequest(ContextCacheRequest const&) = delete#
-
ContextCacheRequest &operator=(ContextCacheRequest const&) = delete#
-
~ContextCacheRequest() = default#
-
std::vector<int32_t> const &prefillStarts() const noexcept#
Per-sequence logical offsets at which runtime prefill begins.
-
int32_t reuseTokenLength(int32_t slot) const noexcept#
Number of reused prefix tokens for a slot (== the logical prefill start offset).
- AdmitSequenceStatus admitSequence(
- std::vector<int32_t> const &tokenIds,
- std::string const &loraWeightsName,
- DecodingKvHeadroom const &headroom,
- int32_t &prefillStart,
- ResidentRef resident,
- cudaStream_t stream,
- std::vector<int32_t> const &mediaTokenIds = {},
- std::vector<imageUtils::ImageData> const &imageBuffers = {},
- std::vector<audioUtils::AudioData> const &audioBuffers = {}
Join one more text-only sequence to this live request: lookup, lease, and row binding. On kAdmitted,
prefillStartreceives the reused prefix length the seated prefill skips.
- bool finalizeSequenceAdmission(
- int32_t slot,
- int32_t const &lookaheadToken,
- int32_t fullInputLength
Record the seated prefill’s lookahead token so the slot’s ledger matches its pages, and publish the ready prefix blocks. Call once, right after the seated prefill succeeds.
-
bool retractSequenceAdmission() noexcept#
Undo the most recent admitSequence before its slot ever joined the runtime batch: the recovery path for a seating that threw between lease and slot append.
- bool publishHybridMtpEndpoint(
- int32_t slot,
- int32_t residentStateLength,
- Tensor const &baseHiddenStates,
- int32_t boundaryHiddenRow
Publish one Hybrid+MTP checkpoint at the stable predecessor boundary. Forwards to the coordinator’s dedicated MTP publication entrypoint; the runtime drives this after the folded draft prefill materialized boundary state.
- bool restoreHybridMtpBoundaryHidden(
- int32_t slot,
- Tensor &baseHiddenStates,
- int32_t destinationRow
Restore the reused checkpoint’s saved boundary base-hidden row into baseHiddenStates for the fold micro-forward.
-
bool preparePrefill()#
-
bool enqueuePrefillCaptures()#
- bool completePrefill(
- DecodingInferenceContext const &context,
- std::vector<int32_t> const &commonStateLengths
- bool prepareDecodeStep(
- DecodingInferenceContext const &context,
- DecodingKvHeadroom const &headroom
- bool completeDecodeStep(
- DecodingInferenceContext const &context,
- std::vector<int32_t> const &commonStateLengths
- bool beginBatchCompaction(
- std::vector<int32_t> const &oldToNew,
- int32_t newBatchSize,
- Tensor &deviceBatchMapping
- bool completeBatchCompaction(
- std::vector<int32_t> const &keepMapping
Compact the per-slot reuse bookkeeping to survivors.
keepMappingis the runtime’s oldSlot -> newSlot batch mapping (-1 for an evicted slot), the same one performBatchEvict uses; without it the reuse-length vector desyncs from the batch after the first eviction.
-
bool finish()#
Public Static Functions
- static std::optional<ContextCacheRequest> begin(
- ContextCacheCoordinator &coordinator,
- LLMGenerationRequest const &request,
- DecodingInferenceContext const &context,
- bool speculativeRequest,
- DecodingKvHeadroom const &headroom,
- std::vector<int32_t> const &mediaTokenIds = {},
- DecodingTokenStateContract tokenStateContract = DecodingTokenStateContract::kCommittedPlusLookahead,
- ContextCacheCommitPolicy commitPolicy = ContextCacheCommitPolicy::kIncludingGeneratedTokens
Admit one tokenized request and bind its cache resources. A disengaged result means admission failed.
- Parameters:
mediaTokenIds – Placeholder token IDs for media modalities (e.g. image, audio). Positions matching any of these IDs are content-hashed for cache differentiation.
- Throws:
std::runtime_error – if a media position has to be hashed and its pixels are not readable on the host.
-
ContextCacheRequest(ContextCacheRequest&&) noexcept = default#