cuda.core.VirtualMemoryResource#
- class cuda.core.VirtualMemoryResource(
- device_id: Device | int,
- config: VirtualMemoryResourceOptions | None = None,
Create a device memory resource that uses the CUDA VMM APIs to allocate memory.
- Parameters:
device_id (Device | int) – Device for which a memory resource is constructed.
config (VirtualMemoryResourceOptions, optional) – A configuration object for the VirtualMemoryResource
Warning
This is a low-level API that is provided only for convenience. Make sure you fully understand how CUDA Virtual Memory Management works before using this. Other MemoryResource subclasses in cuda.core should already meet the common needs.
A shared buffer’s descriptor comes from the exporting process and is not trusted:
Buffer.from_ipc_descriptor()checks what it can, and the driver rejects a size that does not match the exported allocation. Unpickling a shared buffer performs a live import, so unpickle buffers only from a trusted principal. A descriptor pins the physical memory while it exists; release it when the receiver has imported it.Notes
Every buffer this resource returns is a
VirtualMemoryBufferthat owns its address reservations, physical allocations and mappings; closing the buffer releases them.deallocate()is not involved in that path.Virtual memory is unmapped synchronously, so closing a buffer waits for the work queued on its deallocation stream. A buffer released by the garbage collector waits at that point instead. To control when the wait happens, close the buffer explicitly or record an idle stream with
Buffer.set_deallocation_stream().Buffers can be shared with other processes when
config.handle_typeis"posix_fd"(Linux);handle_typeis the switch andis_ipc_enabledthe query.Buffer.ipc_descriptorexports the physical allocations that back a buffer, one file descriptor each, into a new descriptor on every access;Buffer.from_ipc_descriptor()imports a descriptor with aVirtualMemoryResourceof the receiving process for the device that owns the memory; and a buffer sent throughmultiprocessingdoes both. Plainpicklecannot carry the file descriptors; a process with its own file descriptor passing sends the descriptor’sfds,sizes,handle_type, andsizeand rebuilds it withVirtualMemoryIPCBufferDescriptor.from_fds(). The importing resource maps the memory for its device and the devices in itspeersoption, with its own access options. A descriptor pins the physical memory while it exists, in every process that holds one, and nothing else does. A buffer that contains imported memory (an import, or anythingmodify_allocation()derived from one) cannot be exported again; forward the descriptor it came from, or copy into a buffer you own.Methods
- __init__(*args, **kwargs)#
- allocate(
- self,
- size_t size,
- *,
- stream: Stream | GraphBuilder | None = None,
Allocate a buffer of the given size using CUDA virtual memory.
- Parameters:
size (int) – The size in bytes of the buffer to allocate. It is rounded up to the allocation granularity; the returned buffer reports the rounded size.
stream (
Stream|GraphBuilder, optional) – Keyword-only. The allocation itself is synchronous. A real stream is recorded as the buffer’s deallocation stream and synchronized when the buffer closes; with None or a default-stream token the legacy default stream of the resource’s device is recorded instead. A host-located resource records no default stream: its buffers close without a synchronization unless a real stream was given.
- Returns:
A buffer that owns its reservation, physical allocation and mapping.
- Return type:
- Raises:
CUDAError – If any CUDA driver API call fails during allocation. Nothing is left allocated when this method raises.
OverflowError – If
sizerounded up to the granularity does not fit insize_t.
- deallocate(
- self,
- ptr: DevicePointerType,
- int size: int,
- *,
- stream: Stream | GraphBuilder | None = None,
Unmap and free one address range that was reserved and mapped outside this resource.
Buffers returned by
allocate()andmodify_allocation()free themselves when they close and never call this method. It exists for raw pointers wrapped withBuffer.from_handle()withmrset to this resource: the range must be exactly one reservation, and the caller must already have released its owncuMemCreatehandle, so the physical memory is freed by the unmap.- Parameters:
ptr (DevicePointerType) – The start of the reservation.
size (int) – The size of the reservation in bytes.
stream (
Stream|GraphBuilder, optional) – Keyword-only. If given,stream.sync()is called before the range is unmapped, except for a default-stream token on a host-located resource, which has no context to synchronize in.
- modify_allocation(
- self,
- Buffer buf: Buffer,
- size_t new_size,
- config: VirtualMemoryResourceOptions | None = None,
Grow a buffer of this resource to at least
new_sizebytes.The buffer passed in stays open and usable, and is never returned. The returned buffer aliases it: both map the same physical memory, which is freed when the last of the two closes. When the driver can extend the address range in place, the returned buffer has the same pointer; otherwise it has a new one and the existing contents are reachable through both. Closing the returned buffer never closes
buf. A buffer grown from an imported buffer maps imported memory and cannot be exported again.Concurrent calls on buffers that alias one another are safe.
- Parameters:
buf (VirtualMemoryBuffer) – A buffer returned by
allocate()or by this method.new_size (int) – The requested total size in bytes; rounded up to the granularity.
config (VirtualMemoryResourceOptions, optional) – Configuration for the new physical memory chunk only. Existing chunks keep the access they were created with, and the resource’s own configuration is unchanged. It must name the resource’s
location_typeandhandle_typeand passes the same checks as the constructor. Whenbufalready coversnew_sizethere is no new chunk, soconfighas no effect. This method never changes the access of memory that is already mapped.
- Returns:
A new buffer of at least
new_sizeand at leastbuf.sizebytes. Whenbufalready covers the request, the result is a full alias of it and no driver call is made.- Return type:
- Raises:
TypeError – If
bufdid not come from this resource.ValueError – If
confignames a different location or handle type than the resource, or the constructor would reject it.RuntimeError – If
bufis closed, orconfigrequests GPUDirect RDMA on a device without support.OverflowError – If
new_sizerounded up to the granularity does not fit insize_t.CUDAError – If a driver call fails.
bufis untouched when this method raises.
Attributes
- config#
- device#
- device_id#
Get the device ID associated with this memory resource.
- Returns:
int: CUDA device ID. -1 if the memory resource allocates host memory
- is_device_accessible#
Indicates whether the allocated memory is accessible from the device.
- is_host_accessible#
Indicates whether the allocated memory is accessible from the host.
- is_ipc_enabled#
Whether buffers of this resource can be shared with other processes.
True when
config.handle_typeis"posix_fd"on Linux;handle_typeis the only switch, there is no separate option. Withhandle_type=Nonethe allocations cannot be exported, Win32 KMT handles are not transported by cuda.core, and fabric handles are not supported yet.
- is_managed#
bool
Whether buffers allocated by this resource are CUDA managed (unified) memory.
- Type:
MemoryResource.is_managed