Shared Resources#

struct SharedResources#

Process-lifetime resources shared across runners.

Public Functions

inline KVPageTable *getSwaKVPageTable(
int32_t cacheManagerIndex
) noexcept#

Return the optional SWA page table for a cache manager.

inline KVPageTable const *getSwaKVPageTable(
int32_t cacheManagerIndex
) const noexcept#

Const overload of getSwaKVPageTable().

Public Members

std::vector<std::unique_ptr<HybridCacheManager>> cacheManagers#

One HybridCacheManager per engine (index 0 = base, 1 = draft for SpecDecode). unique_ptr because HybridCacheManager is move-only.

std::vector<std::unique_ptr<KVPageTable>> kvPageTables#

One full-capacity KVPageTable per cache manager, index-aligned with cacheManagers. Every table starts with the legacy identity mapping. Production context reuse replaces rows with leased global page IDs and compacts metadata without moving KV bytes.

std::vector<std::unique_ptr<KVPageTable>> swaKVPageTables#

Optional independent SWA page table per cache manager, index-aligned with cacheManagers. A non-null entry has the full logical sequence width but its own bounded physical ID space. It is uploaded as all-sentinel state; the runtime SWA cache manager populates and rotates live mappings.

RopeCache ropePool#
std::unique_ptr<LoRAManager> loraManager#
std::unique_ptr<ExternalWeightManager> externalWeightManager#
Tensor zeroBuffer#
Tensor swaKVCacheMode#

Stable one-byte backing storage for the shape-only SWA runtime mode input.

Public Static Functions

static std::unique_ptr<SharedResources> createForLLM(
LLMEngineConfig const &cfg,
std::unordered_map<std::string, std::string> const &loraWeightsMap,
cudaStream_t stream
)#

Build SharedResources for the vanilla single-engine LLM runtime (KV cache, RoPE pool, LoRA manager, external weight manager, zero buffer).

Recurrent / conv state dtypes for hybrid models are read from cfg.recurrentStateDtype / cfg.convStateDtype — they are parsed strictly from config.json by parseEngineConfig.

The returned externalWeightManager is constructed by this factory; the runtime is responsible for loading files, validating against the base engine, and publishing it to a TensorMap. This keeps SharedResources decoupled from EngineExecutor (no engine I/O or validation happens inside this factory).

static std::unique_ptr<SharedResources> createForSpecDecode(
DeploymentConfig const &bundle,
int32_t maxRuntimeBatchSize,
std::unordered_map<std::string, std::string> const &loraWeightsMap,
cudaStream_t stream
)#

Build SharedResources for a two-engine speculative-decoding runtime (base + draft KV caches, shared RoPE pool, LoRA manager, external weight manager, zero buffer).

As with createForLLM, the returned externalWeightManager is constructed by this factory, and the runtime must load files, validate against the base engine, and publish it to a TensorMap. External weights currently apply to the base engine only.

void trt_edgellm::rt::allocateZeroBuffer(
SharedResources &res,
int64_t bytes
)#