linear_attention#
Modules
Saved execution policy for GDN training-time numerical emulation. |
|
Explicit token-state and encoded-update replay references for QAT. |
|
Differentiable explicit per-sequence prefill/decode phase handoff. |
|
Differentiable KDA prefill with stable per-channel decay interactions. |
|
Batched differentiable GDN prefill with materialized numerical boundaries. |
|
Shared helpers for linear-attention quantization. |
Numerical policies and differentiable kernels for linear attention.