Zarr Compression Tuning#
Zarr stores are the primary persistence format for atomic simulation data in the toolkit. Configuring compression, chunking, sharding, and read windows correctly can reduce disk usage and improve training-time I/O throughput. This guide covers the configuration options, codec trade-offs, and practical recipes for common workloads.
Quick start#
The simplest way to enable compression is to pass a
ZarrWriteConfig when creating a writer or
sink:
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from nvalchemi.data.datapipes.backends.zarr import AtomicDataZarrWriter
from zarr.codecs import ZstdCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(compressors=(ZstdCodec(level=3),)),
)
writer = AtomicDataZarrWriter("/data/example.zarr", config=config)
For dynamics trajectories, pass the same config to
ZarrData:
from nvalchemi.dynamics.sinks import ZarrData
sink = ZarrData("/tmp/trajectory.zarr", config=config)
Tip
The configuration classes are Pydantic models, and you do not need to
import and construct them manually: you can pass a dict with the
same structure and keys and under the hood they will be validated
against the configuration classes. Using the classes explicitly is
helpful, however, when working with modern IDEs and language servers
as they tell you what arguments are required, defaults, etc.
Configuration hierarchy#
The toolkit organises Zarr arrays into three logical groups:
Group |
Contents |
Default compression |
|---|---|---|
|
Pointer arrays ( |
None |
|
Positions, forces, energy, atomic numbers, cell, pbc |
None |
|
User-added arrays via |
None |
ZarrWriteConfig lets you set different
ZarrArrayConfig for each group:
config = ZarrWriteConfig(
meta=ZarrArrayConfig(...), # metadata arrays
core=ZarrArrayConfig(...), # core physics arrays
custom=ZarrArrayConfig(...), # user-added arrays
)
Field overrides#
For fine-grained control, field_overrides takes precedence over group defaults.
Resolution order:
field_overrides["positions"] → if present, use this
↓ (not found)
core (group default) → if present, use this
↓ (not configured)
no compression (Zarr defaults)
Tip
Use field_overrides when a single array has different access patterns from
its group — for example, if positions need fast random access while other core
arrays are read sequentially.
Codec comparison#
Zarr v3 supports pluggable codecs via the zarr.abc.codec.Codec interface. The
toolkit writer accepts any codec supported by Zarr; the benchmark CLI exposes the
common choices zstd, lz4, and blosc-zstd.
Codec |
Class |
Strengths |
Weaknesses |
Typical use |
|---|---|---|---|---|
Zstd |
|
Good ratio, fast decompress |
Moderate compress speed |
General purpose, sequential data |
Blosc/LZ4 |
|
Very fast compress+decompress |
Lower ratio |
Trajectories, real-time I/O |
Blosc/Zstd |
|
Blosc blocking + Zstd ratio |
Slightly more complex |
Large arrays, balanced ratio/speed |
Note
Compression level controls the ratio/speed trade-off. Higher levels yield better compression but slower writes. For Zstd, level 3 is a good default; level 5–9 improves ratio modestly at the cost of write throughput. For LZ4, the level parameter has minimal effect—speed is consistently high.
Blosc options#
BloscCodec exposes codec name, compression level, shuffle, and blocksize through
its constructor. Keep these settings explicit in ZarrWriteConfig when you want
reproducible stores:
from zarr.codecs import BloscCodec
compressor = BloscCodec(cname="zstd", clevel=5)
Chunk size tuning#
The chunk_size parameter in ZarrArrayConfig
controls the chunk length along dimension 0 of the stored array. Other
dimensions use the full extent. Because atom-level fields (positions, forces,
atomic_numbers) are stored concatenated along the atom axis — not per
structure — dimension 0 is the total-atoms axis, not the number of structures.
Target chunk size#
The Zarr documentation recommends chunks of at least 1 MB uncompressed for good throughput, particularly when using Blosc. Smaller chunks increase per-chunk overhead (metadata, system calls, compression dictionary resets). Larger chunks reduce the number of I/O operations for sequential reads but increase read amplification for random access — reading a single 50-atom structure (600 bytes of positions) from a 1 MB chunk wastes 99.9 % of the decompressed data.
Access pattern |
Recommended chunk target |
Rationale |
|---|---|---|
Sequential DataLoader |
1–4 MB |
Amortises overhead across many samples |
Trajectory capture (append, then sequential read) |
1 MB |
Balances write latency and read throughput |
Random access (visualisation, single-sample lookup) |
64–256 KB |
Limits read amplification |
Note
Zarr v3 supports sharding, which decouples the read unit (chunk) from the
storage unit (shard). With sharding you can have small chunks for fine-grained
random access grouped into large shards for filesystem efficiency. Set
shard_size on ZarrArrayConfig to
enable it — the shard size must be a multiple of the chunk size.
Back-of-the-envelope formula#
For a stored array whose rows have trailing_dims trailing dimensions and
dtype size d bytes:
The following table gives concrete values for common arrays:
Array |
Trailing dims |
Dtype |
Bytes/row |
chunk_size (1 MB) |
chunk_size (4 MB) |
|---|---|---|---|---|---|
positions |
3 |
float32 |
12 |
83,333 |
333,333 |
forces |
3 |
float32 |
12 |
83,333 |
333,333 |
atomic_numbers |
1 |
int64 |
8 |
125,000 |
500,000 |
energy |
1 |
float64 |
8 |
125,000 |
500,000 |
cell |
9 |
float32 |
36 |
27,778 |
111,111 |
neighbor_list |
2 |
int64 |
16 |
62,500 |
250,000 |
shifts |
3 |
float32 |
12 |
83,333 |
333,333 |
Positions Example#
Energy Example#
Read amplification#
When reading a single structure by index, the reader fetches the slice
positions[atoms_ptr[i]:atoms_ptr[i+1], :] — typically about 50 rows
(600 bytes for float32 positions). With large chunks, most of the decompressed
data is discarded:
chunk_size |
Chunk bytes (positions) |
Amplification (50-atom read) |
|---|---|---|
333,333 |
4 MB |
6,667× |
83,333 |
1 MB |
1,667× |
10,000 |
120 KB |
200× |
For purely sequential workloads, amplification does not matter because every row
is consumed. For shuffled training, amplification depends on the effective read
window: the DataLoader fuses prefetch_factor batches into one read_many call,
and the Zarr reader can group indices that share chunk locality. Larger chunks can
still hurt fully random single-sample access, so prefer smaller chunks or field
overrides when interactive lookup or visualization is a primary workload.
Warning
Atom-level fields (positions, forces, atomic_numbers) are stored as
concatenated arrays of shape [V_total, ...] where V_total is the sum of
atoms across all structures. The chunk_size parameter controls the number of
rows in each chunk, not the number of structures. System-level fields
(energy, cell, pbc) have one row per structure, so chunk_size directly equals
the number of structures per chunk.
Storage estimation#
The tables below assume 50 atoms per structure on average with ~200 edges (a typical cutoff-based neighbour list). Edge arrays dominate storage; many workflows recompute edges at load time via neighbour lists and omit them from the store.
Per-array breakdown (100k structures)#
Array |
Shape |
Dtype |
Uncompressed |
|---|---|---|---|
positions |
[5M, 3] |
float32 |
60 MB |
forces |
[5M, 3] |
float32 |
60 MB |
atomic_numbers |
[5M] |
int64 |
40 MB |
energy |
[100k] |
float64 |
0.8 MB |
cell |
[100k, 3, 3] |
float32 |
3.6 MB |
pbc |
[100k, 3] |
bool |
0.3 MB |
stress |
[100k, 3, 3] |
float32 |
3.6 MB |
virial |
[100k, 3, 3] |
float32 |
3.6 MB |
dipole |
[100k, 3] |
float32 |
1.2 MB |
neighbor_list |
[20M, 2] |
int64 |
320 MB |
shifts |
[20M, 3] |
float32 |
240 MB |
metadata (ptrs, masks) |
— |
mixed |
27 MB |
Total (with edges) |
760 MB |
||
Total (without edges) |
200 MB |
Scaling by dataset size#
Component |
100k |
1M |
10M |
|---|---|---|---|
Node + system core |
173 MB |
1.7 GB |
17 GB |
Edge arrays |
560 MB |
5.6 GB |
56 GB |
Metadata |
27 MB |
267 MB |
2.7 GB |
Total (with edges) |
760 MB |
7.6 GB |
76 GB |
Total (without edges) |
200 MB |
2.0 GB |
20 GB |
With compression#
Codec |
Typical ratio |
100k |
1M |
10M |
|---|---|---|---|---|
Zstd (level 3) |
2–4× |
190–380 MB |
1.9–3.8 GB |
19–38 GB |
LZ4 |
1.5–2.5× |
300–510 MB |
3.0–5.1 GB |
30–51 GB |
Note
Actual ratios depend heavily on data characteristics. Smooth MD trajectories (correlated frames) compress 4–6×; random equilibrium structures compress 2–3×. Integer arrays (atomic numbers, pointers) often compress 5–10× due to repetition. The estimates above include edge arrays; without edges, divide by ~3.8.
The I/O benchmark tool uses purely random tensors, so its measured ratios (~1.75× Zstd, ~1.63× LZ4) represent a worst case. Real molecular data will compress significantly better.
File count#
Without sharding, each chunk becomes a separate file on local stores. A
Zarr store also contains one zarr.json metadata file per array and per
group, so the total file count across the whole store is the sum of
chunk files for every array plus metadata files (~20 for a typical store).
The table below shows chunk files per array for the positions array
([V_total, 3] float32), which is representative of other atom-level arrays:
chunk_size |
100k (V = 5M) |
1M (V = 50M) |
10M (V = 500M) |
|---|---|---|---|
83,333 (1 MB) |
61 |
601 |
6,001 |
10,000 (120 KB) |
500 |
5,000 |
50,000 |
A typical store has ~10 chunked arrays, so multiply by ~10 for total
chunk files, then add ~20 metadata files. At 100k systems with
chunk_size=10,000, the TUI reports ~4,500 total files; at 100k with
chunk_size=83,333, it reports ~690 total files.
With sharding (shard_size=500,000, chunk_size=10,000), the same
100k-system store drops to ~160 total files — a 28× reduction — because
each shard file bundles 50 chunks.
Filesystem metadata overhead becomes significant above ~10,000 files per
array. If you need small chunks for random access at scale, enable sharding
with shard_size or use a cloud object store (S3, GCS via FsspecStore).
Recipes#
Recipe 1: Sequential dataset (best compression)#
Prioritise disk space over write speed. Use Zstd at a moderate level with large chunks (~1 MB per chunk) for sequential reads.
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from nvalchemi.data.datapipes.backends.zarr import AtomicDataZarrWriter
from zarr.codecs import ZstdCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(
compressors=(ZstdCodec(level=5),),
chunk_size=100_000, # ~1.2 MB chunks for positions [V,3] f32
),
)
writer = AtomicDataZarrWriter("/data/example.zarr", config=config)
Recipe 2: Dynamics trajectory (fast I/O)#
Prioritise write throughput for real-time trajectory capture. Use LZ4 with moderate chunks (~120 KB) to balance write latency and random-access readback.
from nvalchemi.dynamics.sinks import ZarrData
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from zarr.codecs import BloscCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(
compressors=(BloscCodec(cname="lz4"),),
chunk_size=10_000, # ~120 KB chunks for positions [V,3] f32
),
)
sink = ZarrData("/tmp/trajectory.zarr", config=config)
Recipe 3: Per-field override (mixed access patterns)#
Use Zstd for most arrays but LZ4 with smaller chunks for positions (frequently accessed for visualisation or neighbour list rebuilds).
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from nvalchemi.data.datapipes.backends.zarr import AtomicDataZarrWriter
from zarr.codecs import ZstdCodec, BloscCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(
compressors=(ZstdCodec(level=3),),
chunk_size=100_000, # 1 MB chunks for sequential core arrays
),
field_overrides={
"positions": ZarrArrayConfig(
compressors=(BloscCodec(cname="lz4"),),
chunk_size=50_000, # ~600 KB: smaller for random access
),
},
)
writer = AtomicDataZarrWriter("/data/mixed.zarr", config=config)
Recipe 4: Sparse data (skip empty chunks)#
For datasets with many optional fields or sparse validity masks, disable writing empty chunks to save space.
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from nvalchemi.data.datapipes.backends.zarr import AtomicDataZarrWriter
from zarr.codecs import ZstdCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(
compressors=(ZstdCodec(level=3),),
write_empty_chunks=False,
),
custom=ZarrArrayConfig(
compressors=(ZstdCodec(level=3),),
write_empty_chunks=False,
),
)
writer = AtomicDataZarrWriter("/data/sparse.zarr", config=config)
Tip
write_empty_chunks=False is especially useful for custom arrays that are only
populated for a subset of structures. Zarr will skip writing chunks that contain
only the fill value, reducing both disk usage and write time.
Recipe 5: Sharded storage (large datasets)#
For datasets with millions of structures, use sharding to keep small read-friendly chunks while reducing the number of storage objects. The shard size must be a multiple of the chunk size.
from nvalchemi.data.datapipes import ZarrWriteConfig, ZarrArrayConfig
from nvalchemi.data.datapipes.backends.zarr import AtomicDataZarrWriter
from zarr.codecs import ZstdCodec
config = ZarrWriteConfig(
core=ZarrArrayConfig(
compressors=(ZstdCodec(level=3),),
chunk_size=10_000, # 120 KB chunks for random access
shard_size=500_000, # 50 chunks per shard, ~6 MB per shard
),
)
writer = AtomicDataZarrWriter("/data/large.zarr", config=config)
Tip
Sharding is particularly valuable on local filesystems with large datasets
where file count can become a bottleneck. With 10M structures and
chunk_size=10,000, you would get 50,000 files per array without sharding
versus only 1,000 shard files with shard_size=500,000.
I/O benchmark tool#
The toolkit ships a command-line benchmark for measuring Zarr write throughput, readback throughput, and compression ratios on synthetic data. Use it to validate storage configuration and readback strategy before committing to a production workflow.
The CLI has two subcommands:
roundtrip— generate synthetic data, write it to a temporary Zarr store, then read it back and report timing.read— benchmark read throughput against a pre-existing Zarr store, without writing anything.
Run nvalchemi-io-test --help to see the available subcommands. Use
roundtrip when you want the benchmark to create a temporary store, and use
read when you already have a representative store on the target filesystem.
Running the roundtrip benchmark#
# Install (if not already)
$ uv sync
# Basic: compare codec overhead across dataset sizes
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--codec zstd --level 3 \
--chunk-size 83333 --edge-chunk-size 62500
# Compare fast batch readback against one-sample-at-a-time
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--read-mode both --batch-size 64 --prefetch-factor 8
# Model shuffled training reads against compressed stores
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--read-order shuffle --batch-size 64 --prefetch-factor 16
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--read-order block-shuffle --read-order-block-size 8192 \
--batch-size 64 --prefetch-factor 16
# Fast codec with smaller chunks for trajectory-style workloads
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 --codec lz4 \
--chunk-size 10000 --edge-chunk-size 10000
# Larger molecules with edge-specific chunking
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--min-atoms 100 --max-atoms 500 \
--codec zstd --chunk-size 83333 --edge-chunk-size 62500
# With sharding enabled
$ nvalchemi-io-test roundtrip -n 1000 -n 10000 \
--chunk-size 10000 --shard-size 500000 \
--edge-chunk-size 10000 --edge-shard-size 500000
# Write a store to a specific directory for later read benchmarking
$ nvalchemi-io-test roundtrip -n 10000 --codec zstd \
--chunk-size 1024 --shard-size 4096 \
--output-dir /scratch/benchmark_stores/
Roundtrip options:
Option |
Default |
Description |
|---|---|---|
|
1000 10000 100000 |
Dataset sizes to benchmark (repeatable) |
|
10 |
Minimum atoms per structure |
|
100 |
Maximum atoms per structure |
|
— |
Compression codec: |
|
3 |
Compression level |
|
— |
Chunk size for node/system arrays |
|
— |
Shard size for node/system arrays |
|
— |
Chunk size for edge arrays (neighbor_list, shifts) |
|
— |
Shard size for edge arrays |
|
|
Readback path to time: |
|
64 |
Number of samples per emitted DataLoader batch in |
|
16 |
Number of emitted batches to fuse into each backend read in |
|
|
Logical read order: |
|
0 |
Random seed for shuffled read orders |
|
8192 |
Contiguous block size for |
|
|
Request pinned CPU tensors in batch read mode |
|
— |
Persist the written store(s) here instead of a temp directory |
Read-only benchmark#
Use nvalchemi-io-test read to benchmark against an existing Zarr store.
This isolates read performance from generation and write overhead, and lets
you test multiple read configurations against the same store without
rewriting it each time.
# Sequential read baseline
$ nvalchemi-io-test read /path/to/store.zarr
# Shuffled access at different read windows
$ nvalchemi-io-test read /path/to/store.zarr \
--read-order shuffle --batch-size 64 --prefetch-factor 8
$ nvalchemi-io-test read /path/to/store.zarr \
--read-order shuffle --batch-size 64 --prefetch-factor 64
# Compare batch vs. single-sample under shuffle
$ nvalchemi-io-test read /path/to/store.zarr \
--read-mode both --read-order shuffle
Read options:
Option |
Default |
Description |
|---|---|---|
|
— |
Path to an existing Zarr store (directory) |
|
|
|
|
64 |
Number of samples per emitted DataLoader batch |
|
16 |
Number of emitted batches to fuse into each backend read |
|
|
|
|
0 |
Random seed for shuffled orders |
|
8192 |
Block size for |
|
|
Request pinned CPU tensors in batch read mode |
Tip
The read subcommand measures the public DataLoader read path by default:
batch_size controls emitted batches, and prefetch_factor controls how
many emitted batches are fused into one backend read. Use single mode only
as a one-sample-at-a-time baseline.
Note
Benchmark batch mode uses Dataset(skip_validation=True) to focus on storage
and batching throughput for stores that are already trusted. If your training
pipeline keeps validation enabled, expect lower end-to-end throughput.
Readback mode: batch vs. single sample#
The benchmark reports write time plus a full-store readback. Readback uses the batch path by default:
$ nvalchemi-io-test roundtrip -n 10000 --codec zstd \
--chunk-size 83333
In batch mode the benchmark uses the toolkit
DataLoader with fused prefetch.
The emitted batch size is controlled by --batch-size; the backend read window is
controlled by --batch-size * --prefetch-factor. The Zarr reader then receives
large read_many(...) requests and can coalesce physical I/O across the requested
indices.
Use single mode to time a one-sample-at-a-time access pattern:
$ nvalchemi-io-test roundtrip -n 10000 --read-mode single
Use both to emit one row per read path from the same written store:
$ nvalchemi-io-test roundtrip -n 10000 \
--read-mode both --batch-size 64 --prefetch-factor 8
batch mode should be faster for DataLoader-style workloads because it amortises
Python dispatch, Zarr array indexing, chunk lookup, decompression setup, and
filesystem metadata access over many samples. single mode remains useful as a
baseline for debugging and for estimating the penalty paid by code that reads one
structure at a time.
Read order: sequential vs. shuffled training access#
For compressed Zarr stores, the logical index order can dominate throughput.
Sequential readback gives the Zarr reader mostly contiguous physical positions.
Fully shuffled readback models DataLoader(shuffle=True): each emitted batch can
contain unrelated samples, but fused prefetch still gives the reader a larger
window of indices to sort and group by chunk locality.
Use --read-order shuffle to benchmark that worst-case training pattern:
$ nvalchemi-io-test roundtrip -n 10000 --codec zstd \
--chunk-size 83333 --edge-chunk-size 62500 \
--read-order shuffle
Use --read-order block-shuffle to model one locality-preserving training
order:
$ nvalchemi-io-test roundtrip -n 10000 --codec zstd \
--chunk-size 83333 --edge-chunk-size 62500 \
--read-order block-shuffle --read-order-block-size 8192
block-shuffle splits the index range into contiguous blocks of
--read-order-block-size samples, shuffles the blocks, and leaves the
indices inside each block in sequential order. For example, with 10,000
samples and a block size of 2,000 the reader sees five blocks in random
order, but within each block it reads indices 0–1,999, 2,000–3,999, etc.
sequentially.
This benchmark mode does not correspond to a specific DataLoader API;
it is a synthetic access pattern that helps you measure how much throughput
you recover when read locality is partially preserved. Compare
block-shuffle against shuffle to quantify the cost of fully random
access. In practice, a
SizeAwareSampler with bin-packing
can produce similar locality as a side-effect of grouping similarly-sized
systems.
Note
When --read-mode both is used, the two read paths run back-to-back against the
same freshly written store. This is useful for relative comparisons, but the
second mode may benefit from filesystem cache. For strict cold-cache numbers,
run batch and single in separate invocations with the same benchmark
configuration.
The following output illustrates the expected shape of the result table. Treat numbers as machine- and store-specific; use the CLI on the target filesystem for decisions.
Zarr I/O Roundtrip Benchmark — no compression
Systems Read path Read order Batch Prefetch Read window Write Read I/O/s
──────────────────────────────────────────────────────────────────────────────────────────
10,000 batch shuffle 64 32 2,048 0.54s 3.17s 2,695
Read performance tuning#
The benchmark commands above measure the public read paths: batch mode uses
the toolkit DataLoader with fused prefetch and single mode calls
reader.read(...) once per sample. In production, validation, batching, and
device-transfer overhead can dominate the end-to-end pipeline. This section
covers the knobs that matter most for read throughput, especially under shuffled
access patterns.
End-to-end read pipeline.#
The read window: prefetch_factor#
DataLoader groups
prefetch_factor consecutive batches into a single
prefetch_fused_batches() call.
The reader sees one large read_many(...) request containing up to
prefetch_factor * batch_size indices instead of many small calls, which lets
the Zarr backend coalesce random indices into larger physical reads.
The synchronous counterpart is
load_batches(), which accepts
one or more batch-index lists and returns one
Batch per list. DataLoader uses this same
batch-construction path when prefetch_factor=0; only the async double-buffered
prefetch is disabled. New code should prefer load_batches(...) for explicit
batch reads rather than calling older one-batch helpers directly.
Larger windows amortise per-call Zarr overhead across more samples. For
shuffled training, a prefetch_factor of 16–32 is a good starting point, but
the best value depends on store size, chunking, compression, filesystem, and
whether pinned memory is enabled. Use the benchmark tool below on a
representative store before treating any value as a default for production.
Tip
For sequential access the reader already detects contiguous runs, so
prefetch_factor=2 is enough. Increase it primarily when
read_order=shuffle or read_order=block-shuffle.
Skipping validation: skip_validation#
By default the Dataset
validates every loaded sample through
AtomicData (Pydantic), which adds CPU overhead.
When the backing store is known to
contain well-formed data — for example, stores written by the toolkit’s
own writer — you can bypass this:
dataset = Dataset(reader=reader, device="cuda:0", skip_validation=True)
InMemoryDataset accepts the same flag when materializing a reader-backed
resident batch:
dataset = InMemoryDataset(reader=reader, device="cuda:0", skip_validation=True)
With skip_validation=True the Dataset or InMemoryDataset constructs
Batch objects directly from raw tensor
dictionaries via
from_raw_dicts(), avoiding per-sample
Pydantic overhead. For InMemoryDataset this speeds up initial materialization;
subsequent DataLoader iteration selects graphs from the resident batch.
Warning
skip_validation trusts the store contents. Use it only with stores
produced by
AtomicDataZarrWriter
or stores whose schema you have already validated independently.
How the Zarr reader coalesces random indices#
The public read_many method delegates raw loading to the Zarr reader’s
batch-oriented _load_many_samples hook. That hook applies several
backend-specific optimisations automatically:
Resolve logical indices: requested logical indices are mapped through the active-sample mask, so soft-deleted samples are skipped consistently.
Sort by physical position: requests are ordered by physical sample index so the underlying storage sees monotonic offsets where possible.
Group by chunk locality: samples that share Zarr chunks are grouped into range reads, with an amplification cap to avoid pathological over-reads when indices are very sparse.
Fallback for fragmentation: highly fragmented requests use orthogonal selections instead of many tiny range reads.
These optimisations are transparent: read_many still returns results in the
caller’s original request order.
Starting configurations#
Access pattern |
|
|
Notes |
|---|---|---|---|
Sequential training |
2–4 |
|
Small windows are usually enough because samples are already contiguous. |
Shuffled training (trusted store) |
16–64 |
|
Larger windows give the Zarr reader more indices to coalesce. |
Shuffled training (untrusted store) |
16–64 |
|
Keeps validation enabled, but validation can dominate end-to-end time. |
Block-shuffle (block ≥ chunk) |
2–8 |
|
Preserves some locality while still mixing batches. |
Note
Treat these as starting points, not throughput guarantees. Benchmark with
nvalchemi-io-test read or nvalchemi-io-test roundtrip using the same
read order, batch size, prefetch factor, compression, and storage backend you
expect in training.
See also#
Data pipeline: The Data Loading Pipeline guide covers readers, datasets, and dataloaders.
Dynamics sinks: The Data Sinks guide explains how
ZarrDataintegrates with snapshot hooks.API reference: