Getting Started¶
DAQIRI's first-run path is: build the library, optionally tune the host for maximum performance, then run a benchmark. This page covers the basic startup steps and links to long-form guides when you need platform-specific setup, bare-metal packaging, or the full API reference.
System Requirements¶
DAQIRI can run a plain Linux socket path on modest hardware, but the common Raw Ethernet, RoCE, and GPUDirect paths depend on NVIDIA networking and GPU capabilities. Start with the hardware you plan to exercise.
Hardware¶
| Component | Required for | Requirement |
|---|---|---|
| Linux host | All paths | Linux kernel 5.15+; Ubuntu 22.04 or 24.04 recommended |
| NVIDIA NIC | Raw Ethernet, RoCE, GPUDirect | ConnectX-6 Dx or later. Packet pacing, accurate timed transmission, and hardware reorder require ConnectX-7 or later. |
| NVIDIA GPU | GPUDirect and GPU post-processing | RTX or Data Center GPU. GeForce is not supported. |
| Hugepages | DPDK Raw Ethernet or kind: huge memory regions |
Reserved hugepages on the host or in the container runtime environment. |
Supported platforms include NVIDIA Data Center systems, NVIDIA IGX, NVIDIA DGX
Spark, and x86_64 systems with the NIC/GPU requirements above.
Required Libraries And Tools¶
The container build is the recommended starting point because it bundles the user-space libraries and builds the patched DPDK used by DAQIRI. For bare-metal builds, install the matching packages yourself by following the Bare-Metal CMake Build tutorial.
| Component | Required for | Notes |
|---|---|---|
| CUDA Toolkit 12.2+ | Build and GPU paths | The container currently ships CUDA 13.1. CUDA Toolkit 13.0+ adds the default GB10 (sm_121) build target. |
| CMake 3.20+ and C++ build tools | All source builds | cmake, a C++ compiler, git, pkg-config, Python, and standard build tooling. |
| DPDK | DPDK Raw Ethernet engine | Included in the DAQIRI container and patched for dma-buf GPUDirect, so nvidia-peermem is not required inside the container. |
| libibverbs / librdmacm / mlx5 provider | RoCE and ibverbs Raw Ethernet engine | Needed for roce:// socket endpoints and the pure-DevX ibverbs raw engine. |
| NIC diagnostic utilities | System setup and benchmarks | ibstat, ibv_devinfo, ibdev2netdev, mlnx_perf, mlxconfig, and related tools are strongly recommended. |
| Vendored submodules | All source builds | third_party/yaml-cpp and third_party/spdlog; initialize submodules before configuring. |
Optional Libraries¶
| Component | Enables | Notes |
|---|---|---|
| pybind11 | Python bindings | Only needed with -DDAQIRI_BUILD_PYTHON=ON. |
| cuFile / GDS | CUDA device-memory burst file writes | Only needed with -DDAQIRI_ENABLE_GDS=ON; host-memory writes use POSIX APIs without GDS. |
| AWS SDK for C++ with S3 | Raw packet writes to S3-compatible object stores | Only needed with -DDAQIRI_ENABLE_S3=ON; the container can build this SDK from source. |
| OpenTelemetry C++ | Metrics instrumentation | Only needed with -DDAQIRI_ENABLE_OTEL_METRICS=ON; applications still configure the SDK reader/exporter. |
| libnuma | NUMA-aware ring, pool, and huge-memory placement | Auto-detected. DAQIRI falls back to first-touch placement when absent. |
Build¶
Choose either the container build or a bare-metal CMake build. The container is the recommended first pass because it carries the DAQIRI source build, patched DPDK, CUDA user-space dependencies, and RDMA libraries in one image.
git clone git@github.com:NVIDIA/daqiri.git
cd daqiri
BASE_TARGET=dpdk DAQIRI_ENGINE="dpdk ibverbs" scripts/build-container.sh
IGX Thor ships CUDA 13.0. Match that version instead of the default CUDA 13.1 image:
Use BASE_IMAGE=torch when you need the Torch or TensorRT dependencies:
This selects the base image only; building the opt-in TensorRT example
applications is a separate DAQIRI_BUILD_APPLICATIONS=ON source-build
workflow covered in the TensorRT inference tutorial.
Bare-metal builds are supported, but the full setup depends on the host distribution, DOCA/CUDA repositories, and DPDK install prefix. Follow Bare-Metal CMake Build for the full dependency list, DPDK patch workflow, installation checks, cleanup commands, and troubleshooting.
git clone git@github.com:NVIDIA/daqiri.git
cd daqiri
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DBUILD_SHARED_LIBS=ON -DDAQIRI_BUILD_PYTHON=OFF -DDAQIRI_ENGINE="dpdk ibverbs"
cmake --build build -j
cmake --install build --prefix /opt/daqiri
After installation, CMake consumers link the exported target with
find_package(daqiri REQUIRED) and
target_link_libraries(my_app PRIVATE daqiri::daqiri). Pkg-config
consumers can use pkg-config --cflags --libs daqiri.
Most users can keep the defaults. Change CMake flags when enabling Python bindings, GDS, S3, OpenTelemetry, a smaller engine set, tests, or a GPU architecture not covered by the default build. See CMake options reference.
Tune The System¶
This step is optional, but recommended before collecting performance numbers. The built-in host checks surface common networking, GPUDirect, hugepage, and affinity issues before you spend time debugging benchmark output.
The script reports what it can inspect automatically. Persistent host changes such as NIC link layer, hugepages, BAR1 size, MRRS, CPU isolation, GPU clocks, and programmable flex parsing are covered in System Configuration.
Run A Benchmark¶
DAQIRI benchmarks pair an executable with a YAML configuration. If you have a
cable looped back between NIC ports on the system, start with a closed-loop Raw
Ethernet run after replacing the <angle-bracket> placeholders in the YAML for
your system.
Other smoke tests exist if you do not have a cable loopback, including hardware loopback on supported NICs and software loopback when no NIC is available. For those paths, or for throughput and latency measurements, follow the benchmark guide that matches your stream:
- Benchmarking overview: choose a stream type and engine.
- Raw Ethernet Benchmarking: DPDK or ibverbs
raw packet benchmarks, loopback setup, flow programming, hardware reorder, and
throughput measurement with
mlnx_perf. - Socket and RDMA Benchmarking: UDP/TCP and RoCE examples.
Next Steps¶
Keep these pages nearby as you go deeper:
- Concepts: stream types, engines, endpoint URI schemes, packets, bursts, segments, flows, queues, memory regions, GPUDirect, and zero-copy ownership.
- Configuration YAML Walkthrough: annotated examples and a decision tree for choosing an example config.
- System Configuration: NIC drivers, link layers, GPUDirect, hugepages, CPU isolation, GPU clocks, and performance tuning.
- Benchmarking: choose an engine, then run socket/RDMA or Raw Ethernet benchmarks.
- API Guide: the DAQIRI application lifecycle and configuration-first model. Runtime queues, memory regions, and dynamic RX flows are covered in C++ API Usage and Python API Usage.