Block Hash Kernel#

void trt_edgellm::rt::launchFnv1aHashChunksKernel(
uint8_t const *deviceData,
size_t totalSize,
int32_t numChunks,
int32_t numCtas,
uint64_t *partialsOut,
cudaStream_t stream
)#

Launch a multi-CTA parallel FNV-1a Phase 1 kernel on device data.

The kernel splits the payload into numChunks equal chunks distributed across numCtas CTAs. Each thread hashes its chunk independently and writes the 128-bit partial digest to partialsOut as two consecutive uint64_t (hi, lo) per chunk. The caller must synchronize stream then perform the Phase 2 reduction on CPU.