Batch Evict Kernels#
-
struct KVLayerInfo#
Per-layer KV cache metadata for batched kernel operations.
- void trt_edgellm::kernel::compactTensorBatch(
- rt::Tensor const &src,
- rt::Tensor const &batchMapping,
- rt::Tensor &dst,
- int32_t oldActiveBatch,
- int32_t newActiveBatch,
- cudaStream_t stream
Compact a stable old-to-new batch mapping in place. Surviving rows must retain order, so every mapped index is no greater than its source index and the surviving destination indices are contiguous.
batchMapping[i]is the destination row or -1 for an evicted row. The batch dimension must be dimension zero.
- void trt_edgellm::kernel::compactExecutionTensorBatch(
- rt::Tensor &tensor,
- rt::Tensor const &batchMapping,
- int32_t oldActiveBatch,
- int32_t newActiveBatch,
- cudaStream_t stream
Compact execution-owned sequence blocks while preserving a flattened token-major leading dimension. The input may be batch-major
[B, ...]or entry-padded token-major[B * rowsPerSequence, ...].
- void trt_edgellm::kernel::compactExecutionRowIndices(
- rt::Tensor &indices,
- rt::Tensor const &batchMapping,
- int32_t oldActiveBatch,
- int32_t newActiveBatch,
- int32_t executionRowsPerSequence,
- cudaStream_t stream
Compact per-sequence absolute execution-row indices and rebase them to the compacted sequence slots. Unlike logical positions and resident indices, these values include the old execution-slot row base.