Layer Debugger#
- static int32_t trt_edgellm::rt::originalRow(
- std::vector<int32_t> const &originalIndices,
- int32_t slot
Per-request debug dumper for 4-layer numeric validation.
When both environment variables below are set, the LLM runtime dumps, for every inference round (round 0 = prefill, round r = decode step r), the last-token logits, the per-layer combined KV cache (only the layers named in the env var), the per-sequence valid lengths, and the per-round generated token ids. Everything is buffered across rounds and written as a single safetensors file at request end so the comparison tool can reconcile it against the PyTorch golden.
EDGELLM_DUMP_LOGITS_KVCACHE_LAYERS number of leading decoder layers k -> dumps layers 0..k-1. EDGELLM_DUMP_LOGITS_KVCACHE_DIR output directory for the dump file.
The two variables are XOR-coupled: setting exactly one is an error.
Per-layer tensors are gathered from resident slots into execution order. The comparison tool slices each sequence to its valid length in PyTorch using the dumped context_lengths.
Safetensors layout (single file, all rounds): round_{r}.logits [activeBatch, vocab] (native dtype) round_{r}.layer_{i}.kv [activeBatch, 2, kvHeads, maxSeqLen, headDim] (native dtype) dim 1: 0 = key, 1 = value round_{r}.context_lengths [activeBatch] int32 round_{r}.generated_token_ids [activeBatch] int32
Scope: the base model under vanilla or speculative decoding. Hooked from
runBaseModelPrefill(round 0),VanillaDecoder::decodeStep, and — via !decoder_utils::dumpSpecRound— a speculative decoder’s verification step. A speculative ! round commits a variable number of tokens per sequence, so itscontext_lengthsdiffer ! between rows of the same round; the comparison tool pairs each row with the golden round of ! the same length rather than by round index. ! ! Optionally also drives teacher-forcing: whenEDGELLM_FORCE_TOKENS_FILEis set the ! dumper overrides each step’s sampled token with the golden’s (see applyForcedTokens()), ! so the run follows the golden token-for-token. Only ever active alongside a dump. class LayerDebugger { public: ! Build a dumper from the environment, or return nullptr when disabled. !
! Accumulate one round’s tensors into the in-memory buffer. ! ! Synchronises
streamfirst, so the KV cache and logits are final. !
! Record how much of each sequence was restored from the context cache instead of ! executed. Call once from prefill, before the first dumpRound(). ! ! Indexed by
original request row, which is why it survives batch compaction: prefill runs ! before any sequence can finish, so there the active slot and the original row coincide. ! ! A request that reuses a cached prefix only executes the suffix after it, so the runtime’s ! token list counts fewer tokens than the cache actually holds. Every dumped !context_lengthsadds this back, which is what keeps the dump comparable to a golden ! that prefilled the whole sequence. void setReusedPrefixLengths(std::vector<int32_t> lengths);! Write all buffered rounds to a single safetensors file. !
! Teacher-forcing: overwrite each active sequence’s sampled token with the forced ! one for this step. No-op unless
EDGELLM_FORCE_TOKENS_FILEwas set at construction. ! ! Call afterdumpRoundso the dump still records the model’s own sampled token; the ! forced token (if any) is what the caller then commits. This decouples the numeric ! comparison from greedy argmax stability — a near-tie argmax flip no longer diverges the ! two sides, while the dump still surfaces where the runtime wouldhave diverged. ! Sequences are addressed across the whole run, not per request: the force-tokens file ! holds one line per sequence in request order, so a run that issues several requests (the ! context-reuse validation sends one per shared-prefix prompt) still lines up with a golden ! that batched them all. !
! True when teacher-forcing tokens were supplied. bool hasForcedTokens() const noexcept { return !mForcedTokens.empty(); }
! Teacher-forcing for a speculative round: trim the acceptance so the tokens this ! round commits are the golden’s. ! ! A speculative round commits several tokens at once, so overwriting them the way ! applyForcedTokens() does would leave their KV entries describing the tokens the draft ! actually proposed. Only slots
[0, acceptLength - 1)have a committed cache entry — ! the last accepted token is the bonus token, whose entry is written next round — so on the ! first slotjthat disagrees with the golden the acceptance is trimmed toj + 1, ! which drops that slot’s cache entry, and only then is the token replaced. ! ! Call beforethe KV-cache commit, unlike applyForcedTokens(). !
private: LayerDebugger(std::set<int32_t> layers, std::string dir, std::vector<std::vector<int32_t>> forcedTokens);
! Read per-sequence forced token ids from
EDGELLM_FORCE_TOKENS_FILE(one line per ! sequence, whitespace-separated ids), or an empty vector when the env var is unset. Emits a ! warning when forcing is enabled, since it overrides the model’s own sampled tokens. static std::vector<std::vector<int32_t>> readForcedTokensFromEnv();! Original request row for an active slot, via
context.batchIndexMapping.- Throws:
std::runtime_error – if exactly one of the two env vars is set (XOR ! violation) or the layer spec is empty / malformed. static std::unique_ptr<LayerDebugger> fromEnv();
- Parameters:
cacheManager – Base-model KV cache manager. !
pageTable – Full-capacity base-model KV page table. !
swaPageTable – Independent sparse page table for reduced SWA layers, or nullptr. !
logits – Device logits tensor [activeBatch, vocab]. !
validLengths – Per-sequence valid KV/sequence length this round. !
originalIndices – Execution row -> original request row mapping used for reporting. !
residentRefs – Execution row -> persistent resident slot mapping used to gather state. !
generatedTokenIds – Host int32 [activeBatch] tokens sampled this round ! (may be nullptr to skip). !
activeBatchSize – Number of active sequences this round. !
stream –
CUDA stream. void dumpRound(HybridCacheManager& cacheManager, KVPageTable const& pageTable, KVPageTable const* swaPageTable,
Tensor const& logits, std::vector<int32_t> const& validLengths, std::vector<int32_t> const& originalIndices,
std::vector<ResidentRef> const& residentRefs, int32_t const* generatedTokenIds, int32_t activeBatchSize,
cudaStream_t stream);
stream – CUDA stream (forwarded to the safetensors writer). void flush(cudaStream_t stream);
genLengths – Per-sequence count of tokens generated so far (== the index to force). !
originalIndices –
context.batchIndexMapping; see dumpRound(). !tokenIds – Host array [activeBatchSize] of sampled tokens, overwritten in place. !
activeBatchSize –
Number of active sequences. void applyForcedTokens(std::vector<int32_t> const& genLengths, std::vector<int32_t> const& originalIndices,
int32_t* tokenIds, int32_t activeBatchSize);
genLengths – Per-sequence count of tokens generated so far. !
originalIndices –
context.batchIndexMapping; see dumpRound(). !acceptLengths – Host [activeBatchSize] accept lengths, trimmed in place. !
acceptedTokenIds – Host [activeBatchSize, maxAcceptDepth] tokens, overwritten in place. !
ownTokens – Out: each sequence’s own token at the slot that ends up last, i.e. ! what it would have committed there without forcing. !
activeBatchSize – Number of active sequences. !
maxAcceptDepth – Row stride of
acceptedTokenIds. !
- Returns:
A dumper if both env vars are set; nullptr if neither is set. !
- Returns:
true if any sequence was trimmed (the caller must then push the arrays back). bool applyForcedAcceptance(std::vector<int32_t> const& genLengths, std::vector<int32_t> const& originalIndices,
int32_t* acceptLengths, int32_t* acceptedTokenIds, std::vector<int32_t>& ownTokens, int32_t activeBatchSize,
int32_t maxAcceptDepth);
-
int32_t trt_edgellm::rt::forcedRowBase(int32_t activeBatchSize)#
This request’s first row in the force-tokens file, claimed on first use.
First use is always prefill, where the active batch is still the full request, so the claim covers every one of its sequences even if some finish later.