Configuration#
Warp has settings at the global, module, and kernel level that can be used to fine-tune the compilation and verbosity
of Warp programs. In cases in which a setting can be changed at multiple levels (e.g.: enable_backward),
the setting at the more-specific scope takes precedence.
Global Settings#
Settings can be modified by direct assignment before or after calling wp.init(),
though some settings only take effect if set prior to initialization.
For example, the location of the user kernel cache can be changed with:
import os
import warp as wp
example_dir = os.path.dirname(os.path.realpath(__file__))
# set default cache directory before wp.init()
wp.config.kernel_cache_dir = os.path.join(example_dir, "tmp", "warpcache1")
wp.init()
See warp.config for a complete list of global settings.
Module Settings#
Module-level settings to control runtime compilation and code generation may be changed by passing a dictionary of
option pairs to wp.set_module_options().
For example, compilation of backward passes for the kernel in an entire module can be disabled with:
wp.set_module_options({"enable_backward": False})
The options for a module can also be queried using wp.get_module_options().
Field |
Type |
Default Value |
Description |
|---|---|---|---|
|
String |
|
A module-level override of the |
|
Integer |
|
A module-level override of the |
|
Integer |
Global setting |
A module-level override of the |
|
Boolean |
Global setting |
A module-level override of the |
|
Boolean |
|
If |
|
Boolean |
|
If |
|
Boolean |
Global setting |
A module-level override of the |
|
Boolean |
Global setting |
A module-level override of the |
|
String |
|
A module-level override of the |
|
Integer |
256 |
The number of CUDA threads per block that kernels in the module will be compiled for. |
|
Boolean |
|
If |
|
Object |
|
Experimental extra build inputs for CPU and CUDA modules. Set to an
|
|
Boolean |
|
A module-level override of the |
|
Boolean |
|
A module-level override of the |
Kernel Settings#
Kernel-level settings can be passed as arguments to the @wp.kernel decorator.
Field |
Type |
Default Value |
Description |
|---|---|---|---|
|
String |
|
Sets the kernel key used for registration and native code generation. If
|
|
Boolean |
|
If |
|
Module | |
|
Controls which module the kernel belongs to. If |
|
int | tuple |
|
CUDA |
|
int |
|
CUDA |
|
Boolean |
|
If |
|
dict |
|
A dict of module-level compilation options to apply to the kernel’s
module. Requires |
|
String |
|
Selects the experimental entry-point ABI. |
@wp.kernel(enable_backward=False)
def scale_2(
x: wp.array[float],
y: wp.array[float],
):
y[0] = x[0] ** 2.0
@wp.kernel(module="unique")
def isolated_kernel(a: wp.array[float], b: wp.array[float]):
# This kernel will be registered in a new unique module created
# just for this kernel and its dependent functions and structs
tid = wp.tid()
b[tid] = a[tid] + 1.0
@wp.kernel(launch_bounds=(256, 1))
def bounded_kernel(a: wp.array[float]):
# CUDA __launch_bounds__ will be set to (256, 1)
tid = wp.tid()
a[tid] = a[tid] * 2.0
@wp.kernel(cuda_max_registers=64)
def register_limited_kernel(a: wp.array[float]):
# CUDA __maxnreg__(64) will be set when supported
tid = wp.tid()
a[tid] = a[tid] * 2.0
@wp.kernel(enable_cuda_smem_spilling=True, launch_bounds=256)
def smem_spilling_kernel(a: wp.array[float]):
tid = wp.tid()
a[tid] = a[tid] * 2.0
@wp.kernel(module_options={"fast_math": True}, module="unique")
def fast_kernel(a: wp.array[float], b: wp.array[float]):
# fast_math is applied to this kernel's unique module
tid = wp.tid()
b[tid] = a[tid] + 1.0
CUDA shared-memory register spilling uses otherwise available shared memory to
reduce local-memory spill traffic. Warp evaluates forward and backward entry
points independently and enables the optimization only when the corresponding
entry point requires no dynamic shared memory, including scratch space required
by tile operations and their callees. Explicit launch_bounds are recommended
to keep the compiler’s shared-memory estimate aligned with the intended block
size. See NVIDIA’s shared-memory register spilling guidance
for performance considerations and CUDA limitations.
CUDA Thread Block Clusters#
CUDA Thread Block Clusters group adjacent CTAs into clusters that the hardware co-schedules on a single GPU Processing Cluster (GPC), unlocking features such as distributed shared memory. Clusters are supported on devices with compute capability 9.0 (Hopper) and above.
The cluster size is declared per-kernel via the cluster_dim decorator
argument — a positive int up to 16 (the default 1 means no clustering):
@wp.kernel(cluster_dim=2)
def my_kernel(a: wp.array[float]):
i = wp.tid()
a[i] = a[i] * 2.0
Warp launches CUDA kernels with a 1D hardware grid, so cluster_dim=N emits
CUDA __cluster_dims__(N, 1, 1); multidimensional cluster shapes are not
supported. Values 2..8 are portable and work on every cluster-capable device;
values 9..16 are non-portable and depend on the GPU’s GPC layout (some
integrated parts such as NVIDIA Thor and DGX Spark GB10 cap at 8). Query
warp.get_cuda_max_cluster_dim() before selecting a value above 8 if you
target a range of devices (it returns 1 on devices that cannot form
clusters).
On a genuine sub-cluster device (compute capability < 9.0) and on CPU,
cluster_dim is silently ignored and the kernel runs unclustered, so portable
code can set it freely. Requesting a cluster while compiling for a target below
sm90 on cluster-capable hardware (for example with
warp.config.ptx_target_arch set below 90) is an error, because the cluster
attribute would be dropped from the generated kernel.
Declaring cluster_dim only forms the cluster; using it — distributed
shared memory, cluster barriers, and cluster rank queries — requires native CUDA
code. See Thread Block Clusters and Distributed Shared Memory in the C++/CUDA workflows guide for a
worked distributed-shared-memory example.
Function Settings#
Function-level settings can be passed as arguments to the @wp.func decorator.
Setting |
Type |
Default Value |
Description |
|---|---|---|---|
|
String |
|
Sets the function key used for registration and native code generation.
If |
|
Module | |
|
Controls which module the function belongs to, following the same rules as the equivalent kernel setting. |
|
Boolean | |
|
Whether the function is inlined into its call sites. |
@wp.func(inline=False)
def expensive_helper(x: wp.vec3) -> wp.vec3:
# kept out of line in every kernel that calls it
return wp.normalize(x) * wp.length(x)
@wp.func(inline=True)
def cheap_helper(x: float) -> float:
# inlined even where the backend compiler would not choose to
return x * 2.0
By default, the backend compiler decides whether to inline a function at each call site. Set
inline=False to prevent inlining or inline=True to require it.
Inlining duplicates the function body at each call site. This can increase register pressure and instruction-cache use. Keeping a function out of line avoids that duplication but adds function-call overhead. The best choice depends on the workload, so measure both options.
The hint covers the generated adjoint as well as the forward function, and is lowered per backend, so the same function remains valid for CPU and CUDA.