cuda::experimental::places::make_partition_descriptor#

Overloads#

make_partition_descriptor(true_dims, spec, grid_dims)#

inline cute_partition_descriptor cuda::experimental::places::make_partition_descriptor(
dim4 true_dims,
const ::std::vector<dim_spec> &spec,
dim4 grid_dims
)

Build a partition from a per-dimension specification.

Each entry of spec describes how the corresponding tensor dimension maps onto the grid (“blocked over axis 0”, …). Split dimensions are padded up to divisibility, which is what makes the resulting layout exact (see the file-level documentation). Every grid axis with extent > 1 must be bound by some entry; unbound axes would leave those places idle (replication is not supported) and are rejected at construction time.

Parameters:
  • true_dims – True tensor extents (dimension 0 fastest)

  • spec – One entry per tensor dimension (at most 4)

  • grid_dims – Extents of the grid of places

make_partition_descriptor(shape, spec, grid_dims)#

template<typename Shape, typename = ::cuda::std::enable_if_t<__has_get_data_dims_v<Shape>>>
cute_partition_descriptor cuda::experimental::places::make_partition_descriptor(
const Shape &shape,
const ::std::vector<dim_spec> &spec,
dim4 grid_dims
)

Build a partition descriptor from a shape object.

Overload of make_partition_descriptor() taking any object exposing dim4 get_data_dims() const (such as the shape of a logical data) in place of the explicit true extents. Forwards to the dim4 overload.

Parameters:
  • shape – [in] Shape object providing the true tensor extents

  • spec – [in] One entry per tensor dimension (at most 4)

  • grid_dims – [in] Extents of the grid of places

Returns:

The partition descriptor built from shape.get_data_dims()