KV Cache Manager#

class KVCacheManager#

Per-layer KV cache manager that supports heterogeneous head configurations across layers. Each attention layer gets its own independently-sized page pool with shape [2, numPages_i, kTOKENS_PER_PAGE, numKVHeads_i, headDim_i], where the K/V split is outermost. Full-capacity layers use the full page budget. SWA-capable layers use either the independent bounded budget or the full budget according to Config::useBoundedSwaKVCache. This replaces the monolithic LinearKVCache allocation when layers have different numKVHeads or headDim.

Public Functions

KVCacheManager() noexcept = default#

Default constructor.

KVCacheManager(Config const &config, cudaStream_t stream)#

Construct and initialize per-layer KV cache.

Allocates one device tensor per attention layer. Once allocated, memory won’t be reallocated. Determines whether all layers share the same numKVHeads and headDim (uniform mode).

Parameters:
  • config – Cache configuration with per-layer configs

  • stream – CUDA stream for allocation

Throws:

std::runtime_error – if config is invalid or data type is unsupported

~KVCacheManager() noexcept#

Destructor.

KVCacheManager(KVCacheManager const&) = delete#

Deleted copy constructor to avoid large data copy.

KVCacheManager &operator=(KVCacheManager const&) = delete#

Deleted copy assignment to avoid large data copy.

Returns:

Reference to this

KVCacheManager(KVCacheManager&&) noexcept#

Move constructor.

KVCacheManager &operator=(KVCacheManager&&) noexcept#

Move assignment operator.

Returns:

Reference to this

rt::Tensor &getCombinedKVCache(int32_t attnLayerIdx) noexcept#

Get the combined KV-cache page pool for the given attention layer.

Parameters:

attnLayerIdx – The index of the attention layer.

Returns:

A reference to the tensor with shape [2, numPages(attnLayerIdx), kTOKENS_PER_PAGE, numKVHeads_i, headDim_i].

rt::Tensor const &getCombinedKVCache(
int32_t attnLayerIdx
) const noexcept#
std::pair<rt::Tensor, rt::Tensor> getSeparateKVCache(
int32_t attnLayerIdx
) const noexcept#

Get the K-half and V-half of the given attention layer’s pool as separate tensor views.

Parameters:

attnLayerIdx – The index of the attention layer.

Returns:

Full-mode layers: [maxBatchSize, capPadded, H, D]. Active bounded layers: [numPages_i, P, H, D].

int32_t maxCapPadded() const noexcept#

Get the padded per-slot token capacity.

Returns:

capPadded = ceil(maxSequenceLength / kTOKENS_PER_PAGE) * kTOKENS_PER_PAGE.

int32_t maxCapPadded(int32_t attnLayerIdx) const noexcept#

Get one layer’s bounded active-private reservation span, padded to pages.

Note

For a reduced layer this is admission metadata, not a dense view of its global pool.

int32_t numPages() const noexcept#

Get the total number of pages spanned by the pool.

Returns:

Config::numPages if non-zero, else the minimum active pages (maxBatchSize * maxCapPadded() / kTOKENS_PER_PAGE).

int32_t numPages(int32_t attnLayerIdx) const noexcept#

Get one layer’s actual physical page count.

bool hasReducedKVCache() const noexcept#

Whether this manager contains a physically reduced SWA pool.

std::optional<int32_t> reducedKVCacheCapacity() const noexcept#

The common active bounded SWA capacity W, or nullopt when the manager uses full storage.

void *kPoolPtr(int32_t attnLayerIdx) const noexcept#

Get the K-half page-pool base pointer for the given attention layer, i.e. the base of the combined page pool.

Parameters:

attnLayerIdx – The index of the attention layer.

Returns:

Device pointer to the K pool.

void *vPoolPtr(int32_t attnLayerIdx) const noexcept#

Get the V-half page-pool base pointer for the given attention layer, i.e. kPoolPtr(attnLayerIdx) offset by numPages(attnLayerIdx) * kTOKENS_PER_PAGE * numKVHeads_i * headDim_i elements.

Parameters:

attnLayerIdx – The index of the attention layer.

Returns:

Device pointer to the V pool.

KVLayerConfig const &getLayerConfig(
int32_t attnLayerIdx
) const noexcept#

Get the layer configuration for the given attention layer.

Parameters:

attnLayerIdx – The index of the attention layer.

Returns:

The KVLayerConfig for this layer.

KVLayerStorageMetadata getLayerStorageMetadata(
int32_t attnLayerIdx
) const noexcept#

Describe the physical pool and logical sequence geometry used by one attention layer.

int32_t numLayers() const noexcept#

Get the number of attention layers.

Returns:

Number of attention layers

bool isUniform() const noexcept#

Check if all layers have the same numKVHeads and headDim.

Returns:

True if all layers are uniform

Config const &getConfig() const noexcept#

Get cache configuration.

Returns:

Cache configuration

struct KVLayerConfig#

Per-layer KV head configuration for heterogeneous models.

Public Functions

inline constexpr KVLayerConfig(
int32_t numKVHeads_ = 0,
int32_t headDim_ = 0,
int32_t kvCacheCapacity_ = 0
) noexcept#

Public Members

int32_t numKVHeads = {}#

Number of key-value heads for this layer.

int32_t headDim = {}#

Head dimension for this layer

int32_t kvCacheCapacity = {}#

Exported bounded-cache capability and window metadata. Runtime storage mode determines whether the layer uses bounded or full physical storage.

struct KVLayerStorageMetadata#

Public Members

KVCacheStorageKind kind = {}#
int32_t physicalPages = {}#
int32_t logicalPagesPerSequence = {}#
int32_t numKVHeads = {}#
int32_t headDim = {}#