Speculative Sampling#

void trt_edgellm::kernel::speculativeSparseAccept(
rt::Tensor const &targetSupportProbabilities,
rt::Tensor const &targetSupportIds,
rt::Tensor const &proposalSupportProbabilities,
rt::Tensor const &proposalSupportIds,
rt::Tensor const &proposalTokenIds,
rt::Tensor const &proposalLengths,
rt::Tensor const &acceptUniforms,
rt::Tensor &acceptedTokenIds,
rt::Tensor &acceptLength,
rt::Tensor *acceptedTokenIndices,
cudaStream_t stream,
rt::Tensor const *maxAcceptLengths
)#

Exact rejection sampling over sparse target/proposal supports.

Each support row must contain unique token ids. The residual distribution is evaluated on support(p) union support(q), so callers such as DFlash2 never allocate [B, P, vocab] proposal probabilities.

void trt_edgellm::kernel::speculativeDenseTargetAccept(
rt::Tensor const &targetProbabilities,
rt::Tensor const &proposalSupportProbabilities,
rt::Tensor const &proposalSupportIds,
rt::Tensor const &proposalTokenIds,
rt::Tensor const &proposalLengths,
rt::Tensor const &acceptUniforms,
rt::Tensor &acceptedTokenIds,
rt::Tensor &acceptLength,
rt::Tensor *acceptedTokenIndices,
cudaStream_t stream,
rt::Tensor const *maxAcceptLengths
)#

Verify sparse draft proposals against a dense target distribution without materializing a dense proposal or residual distribution. Proposal support IDs must be unique within each row.

void trt_edgellm::kernel::speculativeNormalizeTopKTopP(
rt::Tensor const &topKValues,
rt::Tensor &probabilities,
float temperature,
float topP,
cudaStream_t stream
)#

Convert descending top-k logits to the exact temperature/top-p sampling distribution on the fixed sparse support. Entries outside the nucleus are zero and the retained prefix is renormalized.