API Reference
Public APIs for shared logging utilities and telemetry constants.
Note
For configuration options, environment variables, examples, and usage guides, see Configuration Reference.
Log Configuration
Configuration is provided by LogConfig,
documented below.
Core Classes
LogManager
Bases:
objectLog manager for large-scale LLM training.
Supports both regular logging and node local temporary logging. When node local temporary logging is enabled (NVRX_NODE_LOCAL_TMPDIR is set), each node logs independently to avoid overwhelming centralized logging systems. Local rank 0 acts as the node aggregator, collecting logs from all ranks on the same node and writing them to a per-node log file.
Fork-safe: Child processes automatically disable aggregation to avoid conflicts. Service-based: Aggregator can run as a separate service for reliable log collection.
Get the infrastructure local rank (from SLURM_LOCALID env var).
Get the infrastructure rank (from SLURM_PROCID env var).
Get the distributed logger instance.
This property provides direct access to the underlying logger, allowing users to use all standard logging methods: - logger.debug(message) - logger.info(message) - logger.warning(message) - logger.error(message) - logger.critical(message)
Check if node local temporary logging is enabled.
Get the workload local rank (from LOCAL_RANK env var).
Get the workload rank (from RANK env var).
LogConfig
Bases:
objectUtility class for log configuration.
Get infrastructure rank with SLURM job array support.
- For SLURM job arrays, calculates rank as:
array_task_id * nnodes_per_array_task + slurm_procid
- Returns:
Infrastructure rank or None if not available
- Return type:
- Return type:
- Return type:
NodeLogAggregator
Core Functions
setup_logger
Setup the distributed logger.
This function configures the standard Python logger “nvrx” with appropriate handlers for distributed logging. It’s safe to call multiple times - if the logger is already configured, it won’t be reconfigured unless force_reset=True.
The expectation is that this function is called once at the start of the program, and then the logger is used throughout the program i.e. its a singleton.
The logger automatically adapts to distributed or regular mode based on whether NVRX_NODE_LOCAL_TMPDIR is set. If set, enables distributed logging with aggregation. If not set, logs go directly to stderr/stdout.
The logger is fork-safe: all ranks use file-based message passing to ensure child processes can log even when they don’t inherit the aggregator thread.
- Parameters:
node_local_tmp_dir – Optional directory path for temporary files. If None, uses NVRX_NODE_LOCAL_TMPDIR env var.
force_reset – If True, force reconfiguration even if logger is already configured. Useful for subprocesses that need fresh logger setup.
node_local_tmp_prefix (str) – Optional prefix for log files (e.g. “ftlauncher”).
log_file (str | None) – Optional path to log file. When specified, logs are written to this file with rank prefixes (like srun -l) instead of console. All processes write to the same file using append mode for safe concurrent writes.
- Returns:
Configured logger instance
- Return type:
Example
# In main script (launcher.py) or training subprocess from nvidia_resiliency_ext.shared_utils.log_manager import setup_logger logger = setup_logger()
# With log file for consolidated logging across all ranks/nodes logger = setup_logger(log_file=”/path/to/base.log”)
# In subprocesses that need fresh logger setup logger = setup_logger(force_reset=True)
# In other modules import logging logger = logging.getLogger(LogConfig.name) logger.info(“Some message”)
Quick Reference
For quick access to configuration options and environment variables, see the Configuration Reference page which contains:
Complete environment variables reference
Configuration examples and integration guides
Best practices and troubleshooting
Performance considerations and filesystem selection