Paged KV Types#
-
namespace trt_edgellm#
Argument bundles for the CuTe DSL FMHA launchers.
Each public run* entry point fills one of these once from its own arguments and members, and the launcher then spreads it across the generated descriptors. One struct per wrapper shape rather than a single superset: a superset would let a field that matters on one path (kvCacheCapacity on the dense path, tokensPerPage on the paged one) sit silently at zero on another.
These carry no generated types, so this header is safe to include from translation units that use the required FMHA-v2 runner, the optional CUTE_DSL_FMHA_BLACKWELL_ENABLED runner, or both. The descriptor-filling machinery itself lives in cuteDslTensorDescriptors.h, which stays free of any FMHA concept.
Helpers for populating the tensor descriptors emitted by the CuTe DSL C exporter.
Every AOT variant exports its own nominally distinct but layout-identical descriptor structs (
fmha_d64_Tensor_q_tensor_tvsfmha_d128_Tensor_q_tensor_t) plus acute_dsl_<variant>_wrapperentry point. There is no umbrella C type, so descriptor types are recovered here from the signature of the wrapper that consumes them: a call site names only the wrapper and the kernel module, and pairing a descriptor with the wrong variant is not expressible.The exporter (
cutlass/cute/export/c_header_generator.py) always names the membersdata,dynamic_shapesanddynamic_strides, but emits each array only when its dynamic mask is non-empty. A rank-1 descriptor therefore has nodynamic_stridesmember at all and needs makeCuSeqLenTensor() rather than the strided builders below.Small-vocabulary specialisations of the DSpark sampling and verification kernels.
They differ in how the top-k set is found. DSpark selects it in
topKpasses over the row, each a block-wide max-scan, and drops to a single-threaded walk when top-p is used without top-k ortopKexceeds its parallel bound; that cost scales withtopK. A residual-VQ codebook row is small enough to stage in shared memory, so these kernels instead run one MSB-radix select plus a bitonic sort over the surviving candidates — a single pass whose cost is independent oftopK.Semantics, tensor layouts, and accept/residual/bonus behaviour are identical to the DSpark entry points named in each declaration, which remain the fallback above kCpSpecMaxVocab.
-
namespace rt#
Typedefs
-
using PageId = int32_t#
Host identity of one K/V page pair within a single physical KV pool.
Functions
-
inline int32_t pagesPerSlot(int32_t maxCapPadded)#
Number of pages a single slot’s padded token capacity spans.
-
inline int32_t computeMaxPagesPerSeq(int32_t maxKVCacheCapacity)#
Maximum number of pages one sequence’s KV capacity spans (the page table’s last dim).
- inline int32_t resolveKvCacheCapacity(
- int32_t kvCacheCapacity,
- int32_t maxKVCacheCapacity
Resolve a per-layer marker. Missing/zero means the engine’s full capacity.
- inline bool isReducedKvCacheCapacity(
- int32_t kvCacheCapacity,
- int32_t maxKVCacheCapacity
Whether a per-layer marker advertises bounded SWA storage capability.
- inline int64_t computeSwaPrivatePagesPerSlot(
- int64_t slidingWindowCapacity,
- int64_t tokensPerPage,
- int64_t replacementPageCount
Active private reservation for one admitted SWA slot. Two bounded retained images permit chunk transitions while stale pages remain live until execution completes, and replacement pages keep decode advancement writable.
- inline int64_t computeMinimumSwaPoolPages(
- int64_t maxBatchSize,
- int64_t slidingWindowCapacity
Minimum total SWA budget needed for all active slots using the engine page geometry.
- inline int64_t computeMinimumKvPoolPages(
- int64_t maxBatchSize,
- int64_t maxKVCacheCapacity
Computes the paged-KV pool’s minimum active pages:
maxBatchSize * ceil(maxKVCacheCapacity / kTOKENS_PER_PAGE). Identity-mapped active slots occupy these first pages. A build may serialize extra retained pages for cross-request reuse, but its engine profile, runtime registry, allocation, and page table must all use that same count. Callers validate positive inputs and validate againstkMAX_KV_POOL_PAGESbefore narrowing to int32.
Variables
-
int32_t kTOKENS_PER_PAGE = {128}#
Page size P for the paged-KV layout (tokens per page).
-
int32_t kUNUSED_PAGE_ENTRY = {-1}#
Sentinel value for an unallocated page-table entry.
-
int64_t kMAX_KV_POOL_PAGES = {std::numeric_limits<int32_t>::max() / 2}#
Largest K-page count for which both the derived V ids and the CuTe FMHA combined K/V page count (
2 * numPages) fit a positive int32.
-
int32_t kMAX_KV_CACHE_CAPACITY = (std::numeric_limits<int32_t>::max() / kTOKENS_PER_PAGE) * kTOKENS_PER_PAGE#
Largest token capacity whose page-aligned padded value still fits int32.
-
int32_t kSWA_BOUNDARY_HEADROOM_PAGES = {1}#
Bounded SWA reservation constants. The retained image includes a page-boundary allowance; chunk transitions may temporarily hold the old and new images until execution completes.
-
int32_t kSWA_REPLACEMENT_PAGES = {1}#
-
using PageId = int32_t#
-
namespace rt#