Pipeline Io#
-
class AsyncHostStagingFence#
Public Functions
-
AsyncHostStagingFence() = default#
-
AsyncHostStagingFence(AsyncHostStagingFence const&) = delete#
- AsyncHostStagingFence &operator=(
- AsyncHostStagingFence const&
-
AsyncHostStagingFence(AsyncHostStagingFence &&other) noexcept#
- AsyncHostStagingFence &operator=(
- AsyncHostStagingFence &&other
-
~AsyncHostStagingFence()#
-
void wait()#
-
void record(cudaStream_t stream)#
-
AsyncHostStagingFence() = default#
-
struct StreamingPrefillBuffers#
Persistent copies of the prefill-time input embeddings and engine hidden_states output, used by streaming consumers that run concurrently with the base model’s decode loop.
The base model’s
inputsEmbedsandoutputHiddenStatestensors are reshaped to{B, 1, H}and overwritten by every decode step. This struct retains the{B, prefillLen, H}view as it stood at the end of prefill, so consumers reading these buffers do not race with decode writes.Public Functions
- void populateFromPrefill(
- Tensor const &liveInputEmbeds,
- Tensor const &liveEngineHiddenStates,
- int32_t batch,
- int32_t prefillLen,
- int32_t hiddenSize,
- int32_t maxBatch,
- int32_t maxSeq,
- cudaStream_t stream
Allocate on first call (sized to the worst case
{maxBatch, maxSeq, hiddenSize}), reshape to the current request’s{batch, prefillLen, hiddenSize}, and copy from the live PipelineIO buffers onstream. Subsequent calls reuse the same allocation. Must be invoked after prefill and before the first decode step on the same stream so the copies precede any overwrite ofoutputHiddenStates.
-
struct PipelineIO#
All tensors flowing through the inference pipeline. POINTER STABILITY INVARIANT: After buildTensorMap() is called, this struct must not be moved, and deepstackEmbeds must not be resized. TensorMap holds Tensor* pointers into these members — any reallocation invalidates them.
Public Functions
- void uploadRaggedMetadata(
- RaggedExecutionBatch const &batch,
- cudaStream_t stream
Stage and asynchronously upload one already-validated ragged batch.
- void uploadStateIndices(
- std::vector<ResidentRef> const *residentRefs,
- int32_t numSequences,
- cudaStream_t stream
Stage and asynchronously upload active sequence-to-resident-slot indices. A null residentRefs uses identity mapping for runtimes without resident-slot indirection.
-
void waitForStepHostStaging()#
Protect reusable pinned step metadata and ragged token snapshots before CPU reuse.
-
void recordStepHostUploads(cudaStream_t stream)#
Public Members
-
Tensor hostSelectTokenIndices#
CPU (pinned, [maxBatch, 1] INT64) — pairs with selectTokenIndices for H2D staging.
-
Tensor outputHiddenStates#
Engine accept-layer output: the Qwen3-Omni Talker’s feed on both pipelines. Bound to
hidden_stateson the vanilla path, and toaccept_hidden_stateson a SpecDecode base, wherehidden_statesis instead the draft’s post-norm feed inbaseHiddenStates. Binding the latter here would hand the Talker a post-final-norm tensor — degraded audio rather than an error.
-
StreamingPrefillBuffers streamingPrefill#
Per-request copies of
inputsEmbeds/outputHiddenStatesthat streaming consumers (e.g. the Qwen3-Omni Talker) read while the base model’s decode loop overwrites the live buffers. Populated byLLMInferenceRuntimeonly when streaming output is enabled for the request; otherwise the buffers stay empty (no allocation cost).
-
Tensor packedAttentionMask#
Packed proposal attention mask, [physical_tokens, divUp(proposalSize, 32)] INT32. Written by proposal/verify input preparation kernels; consumed by the base and draft engines via the
kAttentionMaskbinding.
-
Tensor specDecodePositionIds#
SpecDecode position IDs, [physical_tokens] INT32. Written by proposal/verify input preparation kernels; consumed by the base and draft engines via the
kAttentionPosIdbinding.
-
Tensor executionPhaseMarker#
Shape-only execution-phase carrier. Plugins branch on its extent and never read its payload.
-
Tensor contextSequenceCountCarrier#
Shape-only [N_context] INT32 carrier. Its payload is never initialized or read.
-
Tensor skipSoftmaxScale#
Shape-only runtime skip-softmax override carrier (data never read); bound with shape [S] where S comes from LLMEngineConfig::skipSoftmaxScaleOverride.
Public Static Functions
- static PipelineIO createForLLM(
- LLMEngineConfig const &cfg,
- cudaStream_t stream
Build PipelineIO for the vanilla single-engine LLM runtime (basic I/O tensors, deepstack embeds, MRope cos/sin cache).
- static PipelineIO createForSpecDecode(
- DeploymentConfig const &bundle,
- int32_t maxRuntimeBatchSize,
- cudaStream_t stream,
- bool hasAcceptHiddenOutput,
- bool hasTreeMetadataInputs
Build PipelineIO for a two-engine speculative-decoding runtime (basic I/O, hidden states, deepstack embeds, MRope cos/sin cache).
hasAcceptHiddenOutputmust say whether the base engine actually exposes theaccept_hidden_statesbinding: allocating regardless would makeoutputHiddenStates.isEmpty()stop meaning “nothing will fill this”, and the Talker would be handed uninitialised memory instead of failing.hasTreeMetadataInputssimilarly reflects the base or draft engine ABI. Some engines retain these optional bindings even when the selected runtime policy uses a linear proposal.
- void trt_edgellm::rt::allocateBasicIO(
- PipelineIO &io,
- int32_t maxBatch,
- int32_t vocabSize
- void trt_edgellm::rt::prepareRaggedKVPageTable(
- PipelineIO &io,
- KVPageTable const &pageTable,
- int32_t numSequences,
- cudaStream_t stream
Gather resident page-table rows into the stable active-step binding after state-index upload.
- void trt_edgellm::rt::prepareRaggedSwaKVPageTable(
- PipelineIO &io,
- KVPageTable const &pageTable,
- int32_t numSequences,
- cudaStream_t stream
Gather bounded-SWA resident rows into its stable active-step binding after state-index upload.
- PipelineIO &io,
- SharedResources &resources,
- LLMEngineConfig const &cfg,
- RaggedExecutionBatch const &batch,
- int32_t kvCacheIndex,
- cudaStream_t stream
Upload one validated execution batch and derive every token-major engine binding backed by persistent resources. MRoPE resident rows must already have been initialized or published by the caller.
- void trt_edgellm::rt::allocateDeepstackEmbeds(
- PipelineIO &io,
- int32_t numFeatures,
- int32_t maxBatch,
- int32_t maxSeq,
- int32_t hiddenSize,
- nvinfer1::DataType dtype
- void trt_edgellm::rt::allocateSpecDecodeHiddenStates(
- PipelineIO &io,
- int32_t maxBatch,
- int32_t maxSeq,
- int32_t baseHiddenDim,
- int32_t draftHiddenDim,
- nvinfer1::DataType dtype,
- bool allocateDraftHiddenStates
- void trt_edgellm::rt::allocateMRope(
- PipelineIO &io,
- int32_t residentRows,
- int32_t activeRows,
- int32_t maxKVCacheCapacity,
- int32_t rotaryDim
- void trt_edgellm::rt::prepareTextOnlyMRope(
- PipelineIO &io,
- LLMEngineConfig const &cfg,
- int32_t activeRows,
- cudaStream_t stream
- PipelineIO &io,
- SharedResources &res,
- LLMEngineConfig const &cfg,
- int32_t physicalTokens,
- int32_t numSequences,
- cudaStream_t stream
Gather the current ragged step’s token-aligned RoPE inputs after metadata upload.
- void trt_edgellm::rt::scatterActiveMRopeToResident(
- PipelineIO &io,
- RaggedExecutionBatch const &batch,
- LLMEngineConfig const &cfg,
- cudaStream_t stream
Publish active-row MRoPE preprocessing output into its resident-slot rows.
- TensorMap &map,
- PipelineIO &io,
- SharedResources &res,
- LLMEngineConfig const &cfg,
- int32_t kvCacheIndex
Populate a TensorMap from PipelineIO + SharedResources for engine binding.
This is the critical glue function that wires all allocated tensors into the name-to-pointer map consumed by TensorRegistry::bindAll().
- Parameters:
map – Output map to populate.
io – Pipeline I/O tensors.
res – Shared resources (KV caches, RoPE pool, LoRA, zero buffer).
cfg – Engine configuration.
kvCacheIndex – Index into res.cacheManagers for the target engine.
- TensorMap &map,
- PipelineIO &io,
- SharedResources &res,
- LLMEngineConfig const &cfg,
- int32_t kvCacheIndex
Populate a TensorMap for a DiffusionGemma unified-backbone engine.
This keeps DiffusionGemma-only phase/canvas bindings out of the default autoregressive tensor-map path while still sharing the common KV/RoPE/state bindings with standard LLM engines.
- void trt_edgellm::rt::bindDiffusionUnifiedBackboneTensors(
- TensorMap &map,
- PipelineIO &io,
- Tensor &logits,
- Tensor &canvasIds,
- Tensor &prevSelfConditioningEmbeds,
- Tensor &nextSelfConditioningEmbeds,
- Tensor &selfConditioningTemperature
Rebind DiffusionGemma unified-backbone tensors for the current denoise, prefill, or commit step. Self-conditioning feedback is hidden-size state ping-ponged by the block-diffusion decoder.
- void trt_edgellm::rt::bindDiffusionUnifiedBackboneSelfConditioningTensors( )#
Rebind only the DiffusionGemma self-conditioning tensors that ping-pong between denoise steps. Static unified-backbone bindings are established by bindDiffusionUnifiedBackboneTensors().
- TensorMap &map,
- PipelineIO &io,
- SharedResources &res,
- LLMEngineConfig const &cfg
Populate a TensorMap for a SpecDecode draft engine. Delegates to
buildTensorMapwithkvCacheIndex=1for the common bindings, then patches in draft-engine- specific bindings (base/draft hidden states in+out, packed proposal attention mask, proposal position IDs).Preconditions:
iomust have been constructed viaPipelineIO::createForSpecDecodefor an EAGLE/MTP-style draft path where draftHiddenStatesIn/Out are populated alongside baseHiddenStates, packedAttentionMask, and specDecodePositionIds. DFlash uses its own draft TensorMap.- Parameters:
map – Output map for the draft engine’s bindings.
io – Pipeline I/O (must be the SpecDecode-flavoured one).
res – Shared resources.
cfg – Draft engine configuration.
- TensorMap &map,
- PipelineIO &io,
- SharedResources &res,
- DeploymentConfig const &bundle
Populate a TensorMap for a Gemma4 MTP assistant draft engine.
Unlike EAGLE/MTP draft engines, Gemma4 assistant engines do not own a draft KV cache. Their
past_key_values_*bindings are zero-copy aliases to the base target KV cache selected bydraftCfg.gemma4MTPKVSharingMap.