cuda-bindings 13.4.1 Release notes#

New APIs#

New APIs from CUDA Toolkit 13.4 are now available in cuda-bindings.

New driver API functions:

  • driver.cuDeviceGetFabricClusterUuid()

  • driver.cuDeviceGetCliqueCount()

  • driver.cuDeviceGetCliqueInfo()

  • driver.cuMemGetLocationInfo()

  • driver.cuGraphAddNode_v3()

  • driver.cuGraphNodeSetParams_v2()

  • driver.cuCheckpointOperationComplete()

New runtime API functions:

  • runtime.cudaMemGetLocationInfo()

New cuFile API functions:

  • cufile.readv()

  • cufile.writev()

New NVML API functions:

  • nvml.system_get_cper_v1()

  • nvml.device_get_bbx_time_data_v1()

  • nvml.device_get_accounting_stats_v2()

  • nvml.device_get_remapped_rows_v2()

  • nvml.device_set_adaptive_tgp_mode_v1()

  • nvml.device_get_adaptive_tgp_mode_info_v1()

  • nvml.device_set_memory_limits_v1()

  • nvml.device_get_memory_limits_v1()

  • nvml.device_get_gpu_fabric_info_v4()

  • nvml.device_perf_metrics_get_samples_v1()

  • nvml.device_set_nvlink_bw_mode_async_v1()

  • nvml.device_get_nv_link_telemetry_samples_v1()

  • nvml.event_set_register_gpu_operational_events_v1()

  • nvml.event_set_wait_v3()

  • nvml.event_set_get_context_count_v1()

  • nvml.event_set_get_context_info_v1()

  • nvml.event_set_get_gpu_operational_event_context_legacy_xid_v1()

  • nvml.device_get_bank_remapper_status_v1()

  • nvml.event_set_get_context_data_v1()

Breaking changes#

Members of cuda-bindings classes that represent memory for use by future versions of the CUDA API were incorrectly exposed to Python. These have been removed and are no longer exposed. These members have names like reserved, internal, padding and unused (sometimes with a numeric suffix). The C++ CUDA documentation may sometimes say that these members “must be zeroed”, but that is not required with the cuda-bindings Python interface (it happens automatically).

This affects the following fields:

  • driver: CUDA_ARRAY_MEMORY_REQUIREMENTS_st.reserved, CUDA_ARRAY_SPARSE_PROPERTIES_st.reserved, CUDA_EXTERNAL_MEMORY_BUFFER_DESC_st.reserved, CUDA_EXTERNAL_MEMORY_HANDLE_DESC_st.reserved, CUDA_EXTERNAL_MEMORY_MIPMAPPED_ARRAY_DESC_st.reserved, CUDA_EXTERNAL_SEMAPHORE_HANDLE_DESC_st.reserved, CUDA_EXTERNAL_SEMAPHORE_SIGNAL_PARAMS_st.reserved, CUDA_EXTERNAL_SEMAPHONE_SIGNAL_PARAMS_st.nvSciSync.reserved, CUDA_EXTERNAL_SEMAPHORE_WAIT_PARAMS_st.reserved, CUDA_EXTERNAL_SEMAPHORE_WAIT_PARAMS_st.nvSciSync.reserved, CUDA_MEMCPY3D_st.reserved0, CUDA_MEMCPY3D_st.reserved1, CUDA_MEMCPY_NODE_PARAMS_st.reserved, CUDA_RESOURCE_DESC_st.res.reserved, CUDA_RESOURCE_VIEW_DESC_st.reserved, CUDA_TEXTURE_DESC_st.reserved, CU_DEV_SM_RESOURCE_GROUP_PARAMS_st.reserved, CUarrayMapInfo_st.reserved, CUcheckpointCheckpointArgs_st.reserved, CUcheckpointLockArgs_st.reserved0, CUcheckpointLockArgs_st.reserved1, CUcheckpointRestoreArgs_st.padding0, CUcheckpointRestoreArgs_st.reserved, CUcheckpointUnlockArgs_st.reserved, CUdevWorkqueueResource_st.reserved, CUgraphEdgeData_st.reserved, CUgraphNodeParams_st.reserved0, CUgraphNodeParams_st.reserved1, CUgraphNodeParams_st.reserved2, CUipcEventHandle_st.reserved, CUipcMemHandle_st.reserved, CUmemAllocationProp_st.allocFlags.reserved, CUmemDecompressParams_st.padding, CUmemPoolProps_st.reserved, CUmemPoolPtrExportData_st.reserved

  • nvml: AccountingStats.reserved, VgpuMetadata.reserved, VgpuPgpuMetadata.reserved

  • runtime: cudaArrayMemoryRequirements.reserved, cudaArraySparseProperties.reserved, cudaDevSmResourceGroupParams_st.reserved, cudaDevWorkqueueResource.reserved, cudaDeviceProp.reserved, cudaEglPlaneDesc_st.reserved, cudaExternalMemoryBufferDesc.reserved, cudaExternalMemoryHandleDesc.reserved, cudaExternalMemoryMipmappedArrayDesc.reserved, cudaExternalSemaphoreHandleDesc.reserved, cudaExternalSemaphoreSignalParams.reserved, cudaExternalSemaphoreWaitParams.reserved, cudaFuncAttributes.reserved, cudaGraphEdgeData_st.reserved, cudaGraphNodeParams.reserved0, cudaGraphNodeParams.reserved1, cudaGraphNodeParams.reserved2, cudaIpcEventHandle_st.reserved, cudaIpcMemHandle_st.reserved, cudaMemFabricHandle_st.reserved, cudaMemPoolProps.reserved, cudaMemPoolPtrExportData.reserved, cudaMemcpyNodeParams.reserved, cudaPointerAttributes.reserved, cudaPointerAttributes.unused, cudaResourceViewDesc.reserved

It is no longer possible to get or set these bytes directly. To create a wrapped struct from a buffer (for example, for pickling), see Setting raw bytes of structs.

Bugfixes#

  • Fixed a bug in the wrapping of nvrtcBundledHeadersInfo. (PR #2754)

  • Fixed CUDA_PYTHON_DISABLE_MAJOR_VERSION_WARNING: previously, setting it to "0" (or any other non-empty string) still disabled the warning; it is now parsed as a bool-like value, so "0" correctly leaves the warning enabled. (PR #2581)

  • Fixed a crash in nvml.system_event_set_wait caused by calling resize() on a non-owning SystemEventData_v1._data view. (PR #2690)

  • get_cuda_native_handle no longer misreports a KeyError raised from within a registered getter as an “Unknown type” error. (PR #2551)

  • Fixed cuFile status checking to no longer raise cuFileError spuriously when CUfileError_t.cu_err is set on a non-error path (for example, BAR-size queries on GH200 systems). (PR #2530)

  • Made param_packer.feed() safe under free-threaded Python by moving its internal state initialization to import time. (PR #2417)

Deprecation Notices#

  • Support for using cuda-bindings with Python 3.10 is deprecated and will be removed in a future version. Python 3.10 reaches end of life in October 2026 per the CPython support cycle.

Preview feature#

A new version of the nvrtc API is available as cuda.bindings._v2.nvrtc. The primary improvements are: (1) raising exceptions rather than returning error codes, (2) uses PEP8-compliant naming, and (3) more performance. This API is still experimental and subject to change.

Known issues#

  • Updating from older versions (v12.6.2.post1 and below) via pip install -U cuda-python might not work. Please do a clean re-installation by uninstalling pip uninstall -y cuda-python followed by installing pip install cuda-python.

  • nvml.system_get_process_name on WSL can return incorrect values. To work around this, set the locale to “C” before calling nvml.device_get_compute_running_processes_v3 (which sets the process names) and before calling nvml.system_get_process_name. cuda_core does this automatically, but users of the raw NVML API will need to do this manually.