Cp Spec Kernels#

bool trt_edgellm::kernel::cpSpecSupportsVocab(int32_t vocabSize)#

Whether the CodePredictor speculative path can use these kernels.

void trt_edgellm::kernel::cpSpecTopKTopPProbs(
rt::Tensor const &logits,
rt::Tensor &probabilities,
int32_t rows,
int32_t vocabSize,
float temperature,
int32_t topK,
float topP,
cudaStream_t stream
)#

Temperature + top-k + top-p filtering into dense probability rows.

Semantics match dsparkLogitsToProbabilities: the top-k logits are softmaxed, then truncated at the top-p mass and renormalized by that mass. One CTA sorts one row, so cost is independent of top-k.

void trt_edgellm::kernel::cpSpecSampleRows(
rt::Tensor const &probabilities,
float const *uniforms,
rt::Tensor &tokenIds,
int32_t rows,
int32_t vocabSize,
cudaStream_t stream
)#

Sample one token per row: tokenIds[r] ~ probabilities[r], driven by uniforms[r].

void trt_edgellm::kernel::cpSpecProbabilisticAccept(
rt::Tensor const &targetProbabilities,
rt::Tensor const &draftProbabilities,
rt::Tensor const &draftTokenIds,
rt::Tensor const &proposalLengths,
float const *acceptUniforms,
rt::Tensor &acceptedTokenIds,
rt::Tensor &acceptLength,
int32_t batchSize,
int32_t draftStride,
int32_t verifyProposalLen,
int32_t vocabSize,
cudaStream_t stream
)#

Speculative-sampling verifier over dense target/draft rows.

Drop-in for dsparkProbabilisticAccept with identical tensor layouts and accept/residual/bonus semantics.