Execution and Performance#

Warp performance depends on more than device kernel time. Python submits work, Warp may compile or load native modules, arrays may allocate or move memory, and CUDA devices usually execute their queued operations asynchronously. Separate those costs before deciding what to optimize.

Guides in this section#