Image Utils#

namespace trt_edgellm

Argument bundles for the CuTe DSL FMHA launchers.

Each public run* entry point fills one of these once from its own arguments and members, and the launcher then spreads it across the generated descriptors. One struct per wrapper shape rather than a single superset: a superset would let a field that matters on one path (kvCacheCapacity on the dense path, tokensPerPage on the paged one) sit silently at zero on another.

These carry no generated types, so this header is safe to include from translation units that use the required FMHA-v2 runner, the optional CUTE_DSL_FMHA_BLACKWELL_ENABLED runner, or both. The descriptor-filling machinery itself lives in cuteDslTensorDescriptors.h, which stays free of any FMHA concept.

Helpers for populating the tensor descriptors emitted by the CuTe DSL C exporter.

Every AOT variant exports its own nominally distinct but layout-identical descriptor structs (fmha_d64_Tensor_q_tensor_t vs fmha_d128_Tensor_q_tensor_t) plus a cute_dsl_<variant>_wrapper entry point. There is no umbrella C type, so descriptor types are recovered here from the signature of the wrapper that consumes them: a call site names only the wrapper and the kernel module, and pairing a descriptor with the wrong variant is not expressible.

The exporter (cutlass/cute/export/c_header_generator.py) always names the members data, dynamic_shapes and dynamic_strides, but emits each array only when its dynamic mask is non-empty. A rank-1 descriptor therefore has no dynamic_strides member at all and needs makeCuSeqLenTensor() rather than the strided builders below.

Small-vocabulary specialisations of the DSpark sampling and verification kernels.

They differ in how the top-k set is found. DSpark selects it in topK passes over the row, each a block-wide max-scan, and drops to a single-threaded walk when top-p is used without top-k or topK exceeds its parallel bound; that cost scales with topK. A residual-VQ codebook row is small enough to stage in shared memory, so these kernels instead run one MSB-radix select plus a bitonic sort over the surviving candidates — a single pass whose cost is independent of topK.

Semantics, tensor layouts, and accept/residual/bonus behaviour are identical to the DSpark entry points named in each declaration, which remain the fallback above kCpSpecMaxVocab.

namespace rt
namespace imageUtils#

Enums

enum class ImageFormat : int32_t#

Source pixel format and memory layout of an image buffer.

Describes the input side only; the pipeline always produces packed RGB8.

Values:

enumerator kRGB8 = 0#

Single plane, packed [H, W, 3], pitch-linear.

enumerator kNV12PL#

Two pitch-linear planes: Y8 [H, W], then UV8 interleaved [(H+1)/2, (W+1)/2].

enumerator kNV12BL#

As kNV12PL but hardware block-linear (512 B GOB = 64 B x 8 rows).

enum class ColorStandard : int32_t#

YUV-to-RGB matrix. Required for YUV formats, ignored for kRGB8. kUnspecified is the value-initialised state and is rejected for YUV formats.

Values:

enumerator kUnspecified = 0#
enumerator kBt601#
enumerator kBt709#
enumerator kBt2020#
enum class ColorRange : int32_t#

Luma and chroma excursion. kLimited puts Y in [16, 235], kFull in [0, 255]. kUnspecified is the value-initialised state and is rejected for YUV formats.

Values:

enumerator kUnspecified = 0#
enumerator kLimited#
enumerator kFull#
enum class InterpolationMode#

Interpolation filter for :func:resizeImage.

Values:

enumerator kLINEAR#

Bilinear.

enumerator kBICUBIC#

Catmull-Rom cubic.

class ImageData#
#include <imageUtils.h>

Image data container (image or video frame stack)

Exactly one of buffer and textures carries the pixels: kNV12BL uses textures, every other format uses buffer, read through layout. The decode entry points produce a uint8 RGB tensor shaped [frames, height, width, channels]; wrapped caller-owned storage may instead hold NV12 planes in one flat allocation. channels is always 3.

Public Functions

ImageData() noexcept = default#

Default constructor (creates uninitialized ImageData)

ImageData(rt::Tensor &&data)#

Construct image data.

Parameters:

data – Image tensor with shape [T, H, W, C]. Single-frame still images use T=1.

Throws:

std::runtime_error – if tensor content not UINT8, tensor shape not 4D, or number of channels not 3

unsigned char *data() const noexcept#

Base of the pixels, for consumers that read them as one packed host-RGB run.

Returns:

Pointer to that run, or nullptr unless buffer is host-resident kRGB8.

int64_t frameBytes() const noexcept#

Reach of a single frame from the base of buffer: the last byte layout addresses over every plane, so it folds in ImageLayout::planeOffsetBytes and is a frame stride only where that offset is zero, as it is for kRGB8.

Returns:

Byte count, or 0 for kNV12BL, whose pixels are sampled rather than addressed.

int64_t addressedBytes() const noexcept#

The same reach over every plane and all frames.

Returns:

Byte count, or 0 for kNV12BL, whose pixels are sampled rather than addressed.

ImageData resizedMeta(int64_t newHeight, int64_t newWidth) const#

Metadata-only copy at a new spatial size: keeps channels, frames and fps, replaces height/width, and carries no pixel buffer (the resized pixels live in the caller’s device tensor).

Returns:

ImageData with the new dimensions and this object’s channels, frames and fps.

Public Members

std::shared_ptr<rt::Tensor> buffer#

Pixel storage (UINT8); read it through layout. Null for kNV12BL.

ImageTexturePlanes textures#

Pixel storage for kNV12BL; unset for every other format.

ImageLayout layout#

Format and colour metadata; for buffer, also its plane addressing.

int64_t width = {0}#

Image width.

int64_t height = {0}#

Image height.

int64_t channels = {0}#

Number of channels (e.g., 3 for RGB)

int64_t frames = {1}#

Number of frames (T); the modality is flagged by isVideo.

double fps = {1.0}#

Video sample fps for MRoPE timestamps; ignored unless isVideo.

bool doResize = {true}#

When false, the vision runner skips its internal resize.

bool isVideo = {false}#

Explicit modality: a single-frame video is still a video.

std::vector<double> timestamps#

Optional source timestamps (seconds); empty assumes uniform fps spacing.

struct ImageLayout#
#include <imageUtils.h>

How to read the bytes of ImageData::buffer: plane placement and alignment padding.

Valid bytes per row, excluding the padding pitchBytes may add, for a frame of W x H: kRGB8 plane 0: H rows, 3W. kNV12* plane 0: H rows, W (Y); plane 1: (H+1)/2 rows, 2 * ((W+1)/2) (interleaved CbCr). Frame t of plane p starts at rawPointer() + planeOffsetBytes[p] + t * frameStrideBytes[p]; all planes and frames live in one allocation.

Public Members

ImageFormat format = {ImageFormat::kRGB8}#
ColorStandard colorStandard = {}#

Ignored for kRGB8. Required for YUV formats.

ColorRange colorRange = {}#

Ignored for kRGB8. Required for YUV formats.

std::array<int64_t, 2> planeOffsetBytes = {}#

Plane 1 is CbCr; unused for kRGB8.

std::array<int64_t, 2> pitchBytes = {}#

Row stride, including any alignment padding.

std::array<int64_t, 2> frameStrideBytes = {}#

Ignored when frames == 1.

struct ImageTexturePlanes#
#include <imageUtils.h>

Texture objects over the planes of a block-linear frame the caller imported into CUDA.

Block-linear pixels have no address a kernel can walk, so they are sampled rather than loaded. The caller owns the objects and the arrays behind them. An unset plane is zero.

Public Members

cudaTextureObject_t plane0 = {}#

Luma: one uint8 sample per texel over width x height.

cudaTextureObject_t plane1 = {}#

Chroma: an interleaved CbCr pair over ((w + 1) / 2) x ((h + 1) / 2).