perf#

Utility functions for performance measurement.

Classes

AccumulatingTimer

A timer that accumulates time across multiple calls and works for both CUDA and non-CUDA operations.

Timer

A Timer that can be used as a decorator as well.

Functions

clear_cuda_cache

Clear the CUDA cache.

get_cuda_memory_stats

Get memory usage of specified GPU in Bytes.

get_used_gpu_mem_fraction

Get used GPU memory as a fraction of total memory.

maybe_clear_cuda_cache

Clear the CUDA cache every _EMPTY_CACHE_CHECK_EVERY calls, if there is slack to reclaim.

report_memory

Simple GPU memory report.

class AccumulatingTimer#

Bases: ContextDecorator

A timer that accumulates time across multiple calls and works for both CUDA and non-CUDA operations.

__init__(name='')#

Initialize AccumulatingTimer.

Parameters:
  • name – Name of the timer for reporting

  • use_cuda – Whether to synchronize CUDA before timing

classmethod get_call_count(name)#

Get the number of calls for a timer.

classmethod get_total_time(name)#

Get the total accumulated time for a timer in milliseconds.

classmethod report()#

Report the accumulated times and call counts.

classmethod reset()#

Reset the accumulated times and call counts.

start()#

Start the timer.

Return type:

None

stop()#

End the timer and return the elapsed time in milliseconds.

Return type:

float

class Timer#

Bases: ContextDecorator

A Timer that can be used as a decorator as well.

__init__(name='')#

Initialize Timer.

start()#

Start the timer.

stop()#

End the timer.

Return type:

float

clear_cuda_cache()#

Clear the CUDA cache.

get_cuda_memory_stats(device=None)#

Get memory usage of specified GPU in Bytes.

get_used_gpu_mem_fraction(device='cuda:0')#

Get used GPU memory as a fraction of total memory.

Parameters:

device – Device identifier (default: “cuda:0”)

Returns:

Fraction of GPU memory currently used (0.0 to 1.0).

Returns 0.0 if CUDA is not available.

Return type:

float

maybe_clear_cuda_cache(slack_bytes=4294967296)#

Clear the CUDA cache every _EMPTY_CACHE_CHECK_EVERY calls, if there is slack to reclaim.

empty_cache() syncs the device and hands cached blocks back to the driver, so calling it after every packed weight costs more than it saves. The counter restarts on each check, so the interval is measured from the last one rather than from process start. A caller that makes fewer than _EMPTY_CACHE_CHECK_EVERY calls may not reclaim at all – use clear_cuda_cache() if you need a guaranteed one.

Parameters:

slack_bytes (int)

Return type:

None

report_memory(name='', rank=0, device=None)#

Simple GPU memory report.