Encoder Embedding Cache#

class EncoderEmbeddingCache#

Content-addressed GPU cache for encoder (ViT / audio) output embeddings.

Saves encoded embeddings keyed by a 128-bit hash of the raw media bytes, allowing subsequent requests with identical media to skip the expensive encoder infer() call. Eviction is LRU when the GPU memory budget is exceeded.

Thread safety: single writer on one CUDA stream. Stores and restores must use that stream.

Public Functions

explicit EncoderEmbeddingCache(int64_t maxBudgetBytes)#

Construct with a GPU memory budget in bytes. A budget of 0 disables the cache.

std::optional<std::reference_wrapper<Tensor const>> lookup(
Hash128 key
)#

Look up a cached embedding by content hash. Returns a const reference to the cached tensor on hit, std::nullopt on miss. Updates last-access time on hit.

std::optional<std::reference_wrapper<EncoderEmbeddingCacheEntry const>> lookupEntry(
Hash128 key
)#

Look up the complete encoder state. References remain valid until eviction or clear(). Updates last-access time on hit.

bool tryRestore(
std::vector<Hash128> const &keys,
std::vector<int64_t> const &tokenLengths,
Tensor &outputEmbedding,
std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
cudaStream_t stream
)#

Restore a complete batch only if every entry matches the preprocessed media layout. Incompatible entries are invalidated so the caller can re-encode and cache their replacements.

void store(
Hash128 key,
Tensor const &embedding,
int64_t numTokens,
int64_t hiddenSize,
cudaStream_t stream
)#

Store an embedding under the given content hash by copying from the source tensor. Evicts LRU entries if the budget would be exceeded. The copy is enqueued on stream.

void storeSlice(
Hash128 key,
void const *devicePtr,
int64_t numTokens,
int64_t hiddenSize,
nvinfer1::DataType dtype,
cudaStream_t stream,
std::vector<std::reference_wrapper<Tensor>> const &features = {},
int64_t tokenOffset = 0
)#

Store a slice of an encoder output as a per-media-item cache entry.

Parameters:
  • key – Content hash of the individual media item

  • devicePtr – Pointer into the encoder output buffer at the item’s byte offset

  • numTokens – Number of encoder output tokens for this item

  • hiddenSize – Hidden dimension (columns) of the embedding

  • dtype – Data type of the embedding elements

  • stream – CUDA stream for the device-to-device copy

  • features – Additional [tokens, hidden] outputs, cached and evicted together with the embedding

  • tokenOffset – First source row for this media item in each additional output

void clear()#

Remove all entries and free GPU memory.

inline int64_t usedBytes() const noexcept#

Current GPU bytes used by cached embeddings and auxiliary features.

inline size_t size() const noexcept#

Number of cached entries.

struct EncoderEmbeddingCacheEntry#

Complete per-media encoder outputs keyed by media content hash.

Public Functions

bool canRestore(
Tensor const &outputEmbedding,
std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
int64_t tokenOffset
) const#

Check destination metadata and capacity without changing any output.

void restore(
Tensor &outputEmbedding,
std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
int64_t tokenOffset,
cudaStream_t stream
) const#

Restore this media item’s outputs into a preprocessed request’s encoder buffers.

Parameters:
  • outputEmbedding – Main encoder embedding buffer

  • outputFeatures – Auxiliary features in encoder order

  • tokenOffset – First destination row for this media item

  • stream – CUDA stream used to store and restore cache entries

Public Members

Tensor embedding#

[numTokens, hiddenSize], fp16 or bf16, GPU

std::vector<Tensor> features#

Additional per-token encoder outputs, including deepstack.

int64_t numTokens = {}#

Actual token count (embedding shape[0])

int64_t hiddenSize = {}#
std::chrono::steady_clock::time_point lastAccess#