cuda-bindings 13.4.0 Release notes#

New APIs#

New APIs from CUDA Toolkit 13.4 are now available in cuda-bindings.

New driver API functions:

  • driver.cuDeviceGetFabricClusterUuid()

  • driver.cuDeviceGetCliqueCount()

  • driver.cuDeviceGetCliqueInfo()

  • driver.cuMemGetLocationInfo()

  • driver.cuGraphAddNode_v3()

  • driver.cuGraphNodeSetParams_v2()

  • driver.cuCheckpointOperationComplete()

New runtime API functions:

  • runtime.cudaMemGetLocationInfo()

New cuFile API functions:

  • cufile.readv()

  • cufile.writev()

New NVML API functions:

  • nvml.system_get_cper_v1()

  • nvml.device_get_bbx_time_data_v1()

  • nvml.device_get_accounting_stats_v2()

  • nvml.device_get_remapped_rows_v2()

  • nvml.device_set_adaptive_tgp_mode_v1()

  • nvml.device_get_adaptive_tgp_mode_info_v1()

  • nvml.device_set_memory_limits_v1()

  • nvml.device_get_memory_limits_v1()

  • nvml.device_get_gpu_fabric_info_v4()

  • nvml.device_perf_metrics_get_samples_v1()

  • nvml.device_set_nvlink_bw_mode_async_v1()

  • nvml.device_get_nv_link_telemetry_samples_v1()

  • nvml.event_set_register_gpu_operational_events_v1()

  • nvml.event_set_wait_v3()

  • nvml.event_set_get_context_count_v1()

  • nvml.event_set_get_context_info_v1()

  • nvml.event_set_get_gpu_operational_event_context_legacy_xid_v1()

  • nvml.device_get_bank_remapper_status_v1()

  • nvml.event_set_get_context_data_v1()

Breaking changes#

Members of cuda-bindings classes that represent memory for use by future versions of the CUDA API were incorrectly exposed to Python. These have been removed and are no longer exposed. These members have names like reserved, internal, padding and unused (sometimes with a numeric suffix). The C++ CUDA documentation may sometimes say that these members “must be zeroed”, but that is not required with the cuda-bindings Python interface (it happens automatically).

This affects the following fields:

  • driver: CUDA_ARRAY_MEMORY_REQUIREMENTS_st.reserved, CUDA_ARRAY_SPARSE_PROPERTIES_st.reserved, CUDA_EXTERNAL_MEMORY_BUFFER_DESC_st.reserved, CUDA_EXTERNAL_MEMORY_HANDLE_DESC_st.reserved, CUDA_EXTERNAL_MEMORY_MIPMAPPED_ARRAY_DESC_st.reserved, CUDA_EXTERNAL_SEMAPHORE_HANDLE_DESC_st.reserved, CUDA_EXTERNAL_SEMAPHORE_SIGNAL_PARAMS_st.reserved, CUDA_EXTERNAL_SEMAPHONE_SIGNAL_PARAMS_st.nvSciSync.reserved, CUDA_EXTERNAL_SEMAPHORE_WAIT_PARAMS_st.reserved, CUDA_EXTERNAL_SEMAPHORE_WAIT_PARAMS_st.nvSciSync.reserved, CUDA_MEMCPY3D_st.reserved0, CUDA_MEMCPY3D_st.reserved1, CUDA_MEMCPY_NODE_PARAMS_st.reserved, CUDA_RESOURCE_DESC_st.res.reserved, CUDA_RESOURCE_VIEW_DESC_st.reserved, CUDA_TEXTURE_DESC_st.reserved, CU_DEV_SM_RESOURCE_GROUP_PARAMS_st.reserved, CUarrayMapInfo_st.reserved, CUcheckpointCheckpointArgs_st.reserved, CUcheckpointLockArgs_st.reserved0, CUcheckpointLockArgs_st.reserved1, CUcheckpointRestoreArgs_st.padding0, CUcheckpointRestoreArgs_st.reserved, CUcheckpointUnlockArgs_st.reserved, CUdevWorkqueueResource_st.reserved, CUgraphEdgeData_st.reserved, CUgraphNodeParams_st.reserved0, CUgraphNodeParams_st.reserved1, CUgraphNodeParams_st.reserved2, CUipcEventHandle_st.reserved, CUipcMemHandle_st.reserved, CUmemAllocationProp_st.allocFlags.reserved, CUmemDecompressParams_st.padding, CUmemPoolProps_st.reserved, CUmemPoolPtrExportData_st.reserved

  • nvml: AccountingStats.reserved, VgpuMetadata.reserved, VgpuPgpuMetadata.reserved

  • runtime: cudaArrayMemoryRequirements.reserved, cudaArraySparseProperties.reserved, cudaDevSmResourceGroupParams_st.reserved, cudaDevWorkqueueResource.reserved, cudaDeviceProp.reserved, cudaEglPlaneDesc_st.reserved, cudaExternalMemoryBufferDesc.reserved, cudaExternalMemoryHandleDesc.reserved, cudaExternalMemoryMipmappedArrayDesc.reserved, cudaExternalSemaphoreHandleDesc.reserved, cudaExternalSemaphoreSignalParams.reserved, cudaExternalSemaphoreWaitParams.reserved, cudaFuncAttributes.reserved, cudaGraphEdgeData_st.reserved, cudaGraphNodeParams.reserved0, cudaGraphNodeParams.reserved1, cudaGraphNodeParams.reserved2, cudaIpcEventHandle_st.reserved, cudaIpcMemHandle_st.reserved, cudaMemFabricHandle_st.reserved, cudaMemPoolProps.reserved, cudaMemPoolPtrExportData.reserved, cudaMemcpyNodeParams.reserved, cudaPointerAttributes.reserved, cudaPointerAttributes.unused, cudaResourceViewDesc.reserved

This does it is no longer to get or set these bytes directly. To create a wrapped struct from a buffer (for example, for pickling), see Setting raw bytes of structs.

Deprecation Notices#

  • Support for using cuda-bindings with Python 3.10 is deprecated and will be removed in a future version. Python 3.10 reaches end of life in October 2026 per the CPython support cycle.

Preview feature#

A new version of the nvrtc API is available as cuda.bindings._v2.nvrtc. The primary improvements are: (1) raising exceptions rather than returning error codes, (2) uses PEP8-compliant naming, and (3) more performance. This API is still experimental and subject to change.

Known issues#

  • Updating from older versions (v12.6.2.post1 and below) via pip install -U cuda-python might not work. Please do a clean re-installation by uninstalling pip uninstall -y cuda-python followed by installing pip install cuda-python.

  • nvml.system_get_process_name on WSL can return incorrect values. To work around this, set the locale to “C” before calling nvml.device_get_compute_running_processes_v3 (which sets the process names) and before calling nvml.system_get_process_name. cuda_core does this automatically, but users of the raw NVML API will need to do this manually.