bmm2_qdq
NVFP4 operand helpers for the attention BMM matmuls.
P, V, and the signed A-side share the low-level nvfp4_scalar_qdq
primitive, but retain thin operand-specific wrappers because their layouts
and amax reductions differ. P is nonnegative with layout [M, K]; V is
signed with layout [K, N]; the signed A-side (Q of BMM1) is [M, K].
All use block-16 scaling along the BMM contraction axis.
Functions
NVFP4-finalize complete block-16 groups in |
- fake_quant_v_onwrite(v_cache, block_table, v_lo, v_hi, *, max_new_tokens, page_size=16, v_qdq_scale=1.0)
NVFP4-finalize complete block-16 groups in
[v_lo, v_hi)in place.max_new_tokensis host metadata used to size the masked launch grid. The grid covers every group that the largest query chunk can complete without reading device metadata.v_loandv_himust describe aligned, completed block-16 boundaries; their device values are not host-validated.- Parameters:
v_cache (Tensor)
block_table (Tensor)
v_lo (Tensor)
v_hi (Tensor)
max_new_tokens (int)
page_size (int)
v_qdq_scale (float)
- Return type:
None