Top-p (Nucleus) Sampling
Restricts token sampling to the smallest vocabulary prefix whose cumulative
probability mass is >= top_p, then renormalizes and samples from that nucleus.
Combined with temperature scaling, optional top-k pre-filtering, optional
min-p filtering, and model-owned repetition-penalty support.
The CLI example uses the Learning-path selector:
export TRTMC=trtmc
# Source build inside the development container:
# export TRTMC=./build/trtmc
Algorithm
Current implementation status: top-p/top-k/min-p filtering is host-side. The
model-owned samplers request CPU logits (LogitsLocation::HOST), so text
generation copies the logits to CPU before applying nucleus filtering and
sampling. The current implementation does not have an on-device top-p sampling
kernel; the on-device sampler path is
greedy argmax, and the optional torch-multinomial path still builds the filtered
distribution on host before invoking its CUDA multinomial helper.
- Sort logits descending, take top-k candidates
top_k <= 0means no top-k limit- when
top_k <= 1and0 < top_p < 1, top-p uses the full vocabulary
- Apply temperature:
logit / temperature - Softmax over top-k
- Apply min-p if enabled by dropping tokens below
min_p * max_probability - Find nucleus: smallest prefix with cumulative probability >= top_p
- Renormalize filtered probabilities
- Sample using xorshift64 RNG seeded by
seed
Runtimes that support repetition_penalty apply it before candidate sorting:
for every token already present in the request history, a positive logit is
divided by the penalty and a negative logit is multiplied by it. 1.0
disables the processor. The LFM2 runtime implements this Hugging Face
convention.
Falls back to argmax (greedy) when all sampling parameters are default, when
temperature <= 0, or when top_p <= 0. Values above top_p=1.0 are clamped
to disabled top-p behavior.
Usage
CLI
$TRTMC run bundle.bundle --runtime-root /opt/trtmc/lib \
--prompt "Once upon a time" \
--temperature 0.7 --top-p 0.9 --min-p 0.05 --top-k 50 \
--repetition-penalty 1.05 --seed 42
Defaults
temperature: 1.0 (no scaling)top_p: 1.0 (disabled -- use all top_k tokens)min_p: 0.0 (disabled)repetition_penalty: 1.0 (disabled)top_k: 1 (greedy argmax unlesstop_pis active)seed: -1 (use deterministic default seed 42)
--top-p 0.9 is meaningful without also passing --top-k: the sampler uses the
full vocabulary as the top-p candidate set. Pass --top-k N to apply top-k first
and then top-p/min-p inside those candidates.
Testing
Unit tests live with each implementing family, for example
families/qwen/tests/cpp/test_qwen_sampler.cpp.
test_top_p_nucleus_restricts_candidates-- with logits 0.1 and top_p=0.7, verifies only token 0 (the nucleus) is ever sampledtest_top_p_disabled_allows_full_topk-- top_p=1.0 allows all candidatestest_combined_top_k_top_p-- combined filteringtest_top_p_alone_uses_full_vocab-- verifiestop_pis not reduced to greedy whentop_kis left at its default value- edge-case tests cover
top_k=0,top_p=0, invalid/clamped values, and seededreset()reproducibility
E2E test: families/qwen/tests/manifests/qwen3-0.6b-topp.json (trace_id: IT-E2E-TOPP-001)
-- runs Qwen3-0.6B with temperature=0.7/top_p=0.9/top_k=50/seed=42 under
the sampling_top_p invariant contract. The contract verifies that the runtime
CLI forwards the requested sampling flags, produces non-empty sampled text, and
replays the same seeded output on the configured determinism rerun. It does not
compare sampled text to a greedy external reference, because stochastic decoding
is not expected to match greedy Hugging Face output token-for-token.
The E2E test covers the user-visible sampling contract, not device placement. Host-side implementation details are family-owned and covered by the owning sampler tests.