Speculative Sampling#
- void trt_edgellm::kernel::speculativeSparseAccept(
- rt::Tensor const &targetSupportProbabilities,
- rt::Tensor const &targetSupportIds,
- rt::Tensor const &proposalSupportProbabilities,
- rt::Tensor const &proposalSupportIds,
- rt::Tensor const &proposalTokenIds,
- rt::Tensor const &proposalLengths,
- rt::Tensor const &acceptUniforms,
- rt::Tensor &acceptedTokenIds,
- rt::Tensor &acceptLength,
- rt::Tensor *acceptedTokenIndices,
- cudaStream_t stream,
- rt::Tensor const *maxAcceptLengths
Exact rejection sampling over sparse target/proposal supports.
Each support row must contain unique token ids. The residual distribution is evaluated on support(p) union support(q), so callers such as DFlash2 never allocate [B, P, vocab] proposal probabilities.
- void trt_edgellm::kernel::speculativeDenseTargetAccept(
- rt::Tensor const &targetProbabilities,
- rt::Tensor const &proposalSupportProbabilities,
- rt::Tensor const &proposalSupportIds,
- rt::Tensor const &proposalTokenIds,
- rt::Tensor const &proposalLengths,
- rt::Tensor const &acceptUniforms,
- rt::Tensor &acceptedTokenIds,
- rt::Tensor &acceptLength,
- rt::Tensor *acceptedTokenIndices,
- cudaStream_t stream,
- rt::Tensor const *maxAcceptLengths
Verify sparse draft proposals against a dense target distribution without materializing a dense proposal or residual distribution. Proposal support IDs must be unique within each row.
- void trt_edgellm::kernel::speculativeNormalizeTopKTopP(
- rt::Tensor const &topKValues,
- rt::Tensor &probabilities,
- float temperature,
- float topP,
- cudaStream_t stream
Convert descending top-k logits to the exact temperature/top-p sampling distribution on the fixed sparse support. Entries outside the nucleus are zero and the retained prefix is renormalized.