Encoder Embedding Cache#
-
class EncoderEmbeddingCache#
Content-addressed GPU cache for encoder (ViT / audio) output embeddings.
Saves encoded embeddings keyed by a 128-bit hash of the raw media bytes, allowing subsequent requests with identical media to skip the expensive encoder
infer()call. Eviction is LRU when the GPU memory budget is exceeded.Thread safety: single writer on one CUDA stream. Stores and restores must use that stream.
Public Functions
-
explicit EncoderEmbeddingCache(int64_t maxBudgetBytes)#
Construct with a GPU memory budget in bytes. A budget of 0 disables the cache.
- std::optional<std::reference_wrapper<Tensor const>> lookup(
- Hash128 key
Look up a cached embedding by content hash. Returns a const reference to the cached tensor on hit, std::nullopt on miss. Updates last-access time on hit.
- std::optional<std::reference_wrapper<EncoderEmbeddingCacheEntry const>> lookupEntry(
- Hash128 key
Look up the complete encoder state. References remain valid until eviction or clear(). Updates last-access time on hit.
- bool tryRestore(
- std::vector<Hash128> const &keys,
- std::vector<int64_t> const &tokenLengths,
- Tensor &outputEmbedding,
- std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
- cudaStream_t stream
Restore a complete batch only if every entry matches the preprocessed media layout. Incompatible entries are invalidated so the caller can re-encode and cache their replacements.
- void store( )#
Store an embedding under the given content hash by copying from the source tensor. Evicts LRU entries if the budget would be exceeded. The copy is enqueued on
stream.
- void storeSlice(
- Hash128 key,
- void const *devicePtr,
- int64_t numTokens,
- int64_t hiddenSize,
- nvinfer1::DataType dtype,
- cudaStream_t stream,
- std::vector<std::reference_wrapper<Tensor>> const &features = {},
- int64_t tokenOffset = 0
Store a slice of an encoder output as a per-media-item cache entry.
- Parameters:
key – Content hash of the individual media item
devicePtr – Pointer into the encoder output buffer at the item’s byte offset
numTokens – Number of encoder output tokens for this item
hiddenSize – Hidden dimension (columns) of the embedding
dtype – Data type of the embedding elements
stream – CUDA stream for the device-to-device copy
features – Additional [tokens, hidden] outputs, cached and evicted together with the embedding
tokenOffset – First source row for this media item in each additional output
-
void clear()#
Remove all entries and free GPU memory.
-
inline int64_t usedBytes() const noexcept#
Current GPU bytes used by cached embeddings and auxiliary features.
-
inline size_t size() const noexcept#
Number of cached entries.
-
explicit EncoderEmbeddingCache(int64_t maxBudgetBytes)#
-
struct EncoderEmbeddingCacheEntry#
Complete per-media encoder outputs keyed by media content hash.
Public Functions
- bool canRestore(
- Tensor const &outputEmbedding,
- std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
- int64_t tokenOffset
Check destination metadata and capacity without changing any output.
- void restore(
- Tensor &outputEmbedding,
- std::vector<std::reference_wrapper<Tensor>> const &outputFeatures,
- int64_t tokenOffset,
- cudaStream_t stream
Restore this media item’s outputs into a preprocessed request’s encoder buffers.
- Parameters:
outputEmbedding – Main encoder embedding buffer
outputFeatures – Auxiliary features in encoder order
tokenOffset – First destination row for this media item
stream – CUDA stream used to store and restore cache entries