Batch Evict Kernels#

struct KVLayerInfo#

Per-layer KV cache metadata for batched kernel operations.

Public Members

void *data#

Base of this layer’s physical K-then-V page-pool allocation.

int32_t numKVHeads#

Number of KV heads for this layer.

int32_t maxSeqLen#

Active-slot row capacity (capPadded)

void trt_edgellm::kernel::compactTensorBatch(
rt::Tensor const &src,
rt::Tensor const &batchMapping,
rt::Tensor &dst,
int32_t oldActiveBatch,
int32_t newActiveBatch,
cudaStream_t stream
)#

Compact a stable old-to-new batch mapping in place. Surviving rows must retain order, so every mapped index is no greater than its source index and the surviving destination indices are contiguous. batchMapping[i] is the destination row or -1 for an evicted row. The batch dimension must be dimension zero.

void trt_edgellm::kernel::compactExecutionTensorBatch(
rt::Tensor &tensor,
rt::Tensor const &batchMapping,
int32_t oldActiveBatch,
int32_t newActiveBatch,
cudaStream_t stream
)#

Compact execution-owned sequence blocks while preserving a flattened token-major leading dimension. The input may be batch-major [B, ...] or entry-padded token-major [B * rowsPerSequence, ...].

void trt_edgellm::kernel::compactExecutionRowIndices(
rt::Tensor &indices,
rt::Tensor const &batchMapping,
int32_t oldActiveBatch,
int32_t newActiveBatch,
int32_t executionRowsPerSequence,
cudaStream_t stream
)#

Compact per-sequence absolute execution-row indices and rebase them to the compacted sequence slots. Unlike logical positions and resident indices, these values include the old execution-slot row base.