prefill#
Batched differentiable GDN prefill with materialized numerical boundaries.
Functions
Normalize GDN inputs and run exact prefix plus configured suffix recurrence. |
- matmul_gdn(q, k, v, g, beta, *, policy, state_qdq=False, state_format='fp8_e4m3', scale=None, initial_state=None, output_final_state=False, use_qk_l2norm_in_kernel=False, use_gate_in_kernel=False, use_beta_sigmoid_in_kernel=False, allow_neg_eigval=False, A_log=None, dt_bias=None, cu_seqlens=None, cu_seqlens_cpu=None, state_v_first=False, chunk_size=64, cp_context=None, prefill_lengths=None)#
Normalize GDN inputs and run exact prefix plus configured suffix recurrence.
Uses FP32 working arithmetic (FP64 for double inputs), with floating QDQ state. Prefixes use exact chunk algebra; no configurable operand rounding is enabled.
- Parameters:
policy (LinearAttentionConfig)