Request Stable Rng#

Defines

EDGELLM_RNG_HD#
namespace trt_edgellm

Argument bundles for the CuTe DSL FMHA launchers.

Each public run* entry point fills one of these once from its own arguments and members, and the launcher then spreads it across the generated descriptors. One struct per wrapper shape rather than a single superset: a superset would let a field that matters on one path (kvCacheCapacity on the dense path, tokensPerPage on the paged one) sit silently at zero on another.

These carry no generated types, so this header is safe to include from translation units that use the required FMHA-v2 runner, the optional CUTE_DSL_FMHA_BLACKWELL_ENABLED runner, or both. The descriptor-filling machinery itself lives in cuteDslTensorDescriptors.h, which stays free of any FMHA concept.

Helpers for populating the tensor descriptors emitted by the CuTe DSL C exporter.

Every AOT variant exports its own nominally distinct but layout-identical descriptor structs (fmha_d64_Tensor_q_tensor_t vs fmha_d128_Tensor_q_tensor_t) plus a cute_dsl_<variant>_wrapper entry point. There is no umbrella C type, so descriptor types are recovered here from the signature of the wrapper that consumes them: a call site names only the wrapper and the kernel module, and pairing a descriptor with the wrong variant is not expressible.

The exporter (cutlass/cute/export/c_header_generator.py) always names the members data, dynamic_shapes and dynamic_strides, but emits each array only when its dynamic mask is non-empty. A rank-1 descriptor therefore has no dynamic_strides member at all and needs makeCuSeqLenTensor() rather than the strided builders below.

Small-vocabulary specialisations of the DSpark sampling and verification kernels.

They differ in how the top-k set is found. DSpark selects it in topK passes over the row, each a block-wide max-scan, and drops to a single-threaded walk when top-p is used without top-k or topK exceeds its parallel bound; that cost scales with topK. A residual-VQ codebook row is small enough to stage in shared memory, so these kernels instead run one MSB-radix select plus a bitonic sort over the surviving candidates — a single pass whose cost is independent of topK.

Semantics, tensor layouts, and accept/residual/bonus behaviour are identical to the DSpark entry points named in each declaration, which remain the fallback above kCpSpecMaxVocab.

namespace rt

Enums

enum class SpecRandomPurpose : uint64_t#

Values:

enumerator kTarget = 0xD6E8FEB86659FD93ULL#
enumerator kProposal = 0xA0761D6478BD642FULL#
enumerator kAccept = 0xE7037ED1A0B428DBULL#
enumerator kResidual = 0x8EBC6AF09C88C6E3ULL#
enumerator kBonus = 0x589965CC75374CC3ULL#

Functions

EDGELLM_RNG_HD constexpr bool shouldUseRequestStableSampling (bool supportsLosslessSampling, bool hasExplicitSeed) noexcept
EDGELLM_RNG_HD constexpr uint64_t requestStableNextAbsolutePosition (uint64_t fullPromptTokenCount, uint64_t generatedTokenCount) noexcept
EDGELLM_RNG_HD constexpr uint64_t requestStableMix64 (uint64_t value) noexcept
EDGELLM_RNG_HD constexpr uint64_t requestStableRandomBits (uint64_t requestSeed, uint64_t absolutePosition, SpecRandomPurpose purpose, uint64_t lane) noexcept
EDGELLM_RNG_HD constexpr float requestStableUniform (uint64_t requestSeed, uint64_t absolutePosition, SpecRandomPurpose purpose, uint64_t lane) noexcept