Dflash2 Grouped Dynamic Conv#
- cudaError_t trt_edgellm::kernel::launchDFlash2GroupedDynamicConv(
- rt::Tensor const &hidden,
- rt::Tensor const &delta,
- rt::Tensor const &base,
- rt::OptionalInputTensor residual,
- rt::Tensor &output,
- cudaStream_t stream
Apply one side of DFlash2’s per-token dynamic grouped depthwise convolution.
hidden: […, blockSize, hiddenSize] delta: […, blockSize, kernelSize, hiddenSize / groupSize] base: [kernelSize, hiddenSize] residual: optional FP32, same shape as hidden output: same shape as hidden, activation dtype without residual and FP32 with residual
A tap never reads across a logical block boundary. Production block sizes are in [2, 16]; kernelSize=2 and groupSize=16 vectorize two adjacent channels. Explicit K=1..4 shapes share the same single-launch backend; unsupported layouts return cudaErrorInvalidValue rather than falling back to decomposed operations.