mlflow#

Record a script run on an MLflow tracking server.

Lets an example script upload its invocation, configuration, log and outputs so the run can be reproduced from its MLflow entry alone. mlflow is an optional dependency, imported only once tracking is actually enabled.

Classes

MlflowRunLogger

Record one script invocation as an MLflow run.

Functions

add_mlflow_args

Add --mlflow, --mlflow_experiment and --mlflow_run_name to parser.

command_text

The invocation, as a copy-pasteable line, with credentials masked.

current_user

Return the current username, or "unknown" if the uid has no passwd entry.

default_experiment_name

Build an experiment name of the form <user>/<tool>/<model>-<variant>.

drop_experiment_json

Remove a provenance pointer an untracked export would otherwise inherit.

resolve_mlflow_args

Settle where tracking is configured from, and name the experiment, in place.

resolve_tracking_uri

Settle the tracking URI from --mlflow and the environment.

validate_tracking_uri

Validate an MLflow tracking URI and return it without a trailing slash.

class MlflowRunLogger#

Bases: object

Record one script invocation as an MLflow run.

start() opens the run before the expensive work begins, so a bad URI, a missing token or an unreachable server fails there rather than after hours; it also uploads the invocation and any configuration passed to it, which keeps a crashed run useful. finish() uploads the captured log plus any outputs and closes the run. Everything is a no-op when enabled is false, so callers need no branching.

While the run is open, stdout/stderr are teed to a file that is uploaded as logs/<script>.log. Logging handlers that libraries bound to sys.stderr at import time are re-pointed at the tee for the duration and handed back afterwards.

Failures after the run is open are reported as warnings and never raised: losing a tracking server must not turn a successful job into a failed one.

Note

command.txt masks --*token*-style option values and credentials embedded in a URI, but the captured log is whatever the script printed, so a secret echoed to stdout still reaches the server. Prefer passing credentials via the environment.

tracking_uri must already be validated (see validate_tracking_uri()), experiment_name is created if absent, run_name defaults to the UTC start time YYYYmmdd-HHMMSS, and enabled=False makes every method a no-op – which is how callers skip non-main ranks or an absent flag. required=False additionally downgrades a failure to open the run into a warning: use it when tracking was inferred from the environment rather than asked for, so an uninstalled client or an unreachable server cannot take the job down with it.

Example

>>> logger = MlflowRunLogger(uri, "alice/hf_ptq/Qwen3-0.6B-nvfp4")
>>> logger.start(params={"model": ckpt}, texts={"config.yaml": config_yaml})
>>> status = "FAILED"
>>> try:
...     quantize_and_export()
...     status = "FINISHED"
... finally:
...     logger.finish(status, files={"summary/report.txt": report_path})
__init__(tracking_uri, experiment_name, run_name=None, enabled=True, required=True)#

Configure the run without contacting the server; see the class docstring.

Parameters:
  • tracking_uri (str)

  • experiment_name (str)

  • run_name (str | None)

  • enabled (bool)

  • required (bool)

finish(status, texts=None, files=None, metrics=None)#

Upload the run’s outputs and close it with status, e.g. "FINISHED".

texts and files both map artifact path to content, from memory and from disk respectively. A files entry is skipped when its file is absent, or was last modified before the run started – so callers can list optional outputs, and a run that produced none of them does not upload a previous run’s leftovers. metrics merges over the default total_time_s.

Parameters:
  • status (str)

  • texts (dict[str, str] | None)

  • files (Mapping[str, Path | str] | None)

  • metrics (dict[str, float] | None)

Return type:

None

log_experiment_json(checkpoint_dir=None)#

Record which MLflow run produced a checkpoint, on the server and in the checkpoint.

Tags point from a run to the checkpoint it wrote; this is the reverse, so a checkpoint found on disk can be traced back to the run that produced it without searching the server. The artifact goes up for any run that opened, so a failure is traceable from the server side too.

checkpoint_dir also writes the JSON there as EXPERIMENT_JSON. Pass it only once the checkpoint is really on disk, since the file claims authorship of the weights sitting next to it: an output directory existing proves nothing, as it may hold a checkpoint from an earlier attempt whose weights this run never touched. Nothing is recorded at all when the run never opened, which a URI taken from the environment reaches by design.

Parameters:

checkpoint_dir (Path | str | None)

Return type:

None

log_text(artifact_path, text)#

Upload text as an artifact while the run is open, best-effort.

For a value that is only settled midway through the run and is worth having even if the run later crashes – the quantization config a calibration is about to apply, say. start() and finish() cover everything known at the two ends.

Parameters:
  • artifact_path (str)

  • text (str)

Return type:

None

property run_info: dict[str, str]#

Identity of this run on the server, or {} before it is open.

Enough for a consumer holding only this run’s outputs to find it again: run_id is MLflow’s own identifier for the run, a uuid4 hex, unique across experiments. Every field is read back off the run the server returned rather than off what was requested, so a run MLflow resolved differently is reported as it really is.

property run_url: str#

Link to this run in the MLflow UI, or "" before the run is open.

start(params=None, tags=None, texts=None, files=None)#

Open the run: capture output, verify the server, upload the inputs.

params are searchable; tags merge over the defaults (user, hostname, ModelOpt version and commit); texts maps artifact path to content, uploaded here rather than at the end so it survives a crash. files names the outputs the run is expected to produce, so finish() can tell them from files that were already there – pass the same mapping to both.

Opening the run is the readiness check: it is MLflow’s own first request, so it honours the client’s TLS and retry configuration rather than second-guessing it. Set MLFLOW_HTTP_REQUEST_MAX_RETRIES to shorten the wait on a dead host.

Raises:
  • ImportError – If mlflow is not installed and required.

  • Exception – Whatever MLflow raises for an unusable server, if required.

Parameters:
  • params (dict[str, Any] | None)

  • tags (dict[str, Any] | None)

  • texts (dict[str, str] | None)

  • files (Mapping[str, Path | str] | None)

Return type:

None

track(params=None, tags=None, texts=None, files=None, metrics=None)#

Open the run for the duration of the block, closing it with the right status.

Mirrors mlflow.start_run(). files and metrics are uploaded when the block exits; naming the paths upfront is fine because only files this run actually wrote are uploaded (see finish()).

Example

>>> with logger.track(params={"model": ckpt}, files={"summary.txt": report}):
...     quantize_and_export()
Parameters:
  • params (dict[str, Any] | None)

  • tags (dict[str, Any] | None)

  • texts (dict[str, str] | None)

  • files (Mapping[str, Path | str] | None)

  • metrics (dict[str, float] | None)

Return type:

Iterator[MlflowRunLogger]

add_mlflow_args(parser, tool, tracks="Track this run on an MLflow server (e.g. https://<your-mlflow-server>/), uploading the command, the resolved configuration, the run log and the run's summaries.", variant_help='recipe name, or the quantization format')#

Add --mlflow, --mlflow_experiment and --mlflow_run_name to parser.

tool names the script in the default experiment <user>/<tool>/<model>-<variant> (see default_experiment_name()), tracks is the leading description of --mlflow – what this particular script uploads – and variant_help says what the script derives the variant from. Pair with resolve_mlflow_args().

The multi-word flags are registered under both the underscored and the dashed spelling: vLLM’s FlexibleArgumentParser rewrites every --foo_bar on the command line to --foo-bar before matching, so the dashed spelling has to exist for the flag to be reachable there at all, and a user moving between the example scripts should not have to remember which spelling each one took.

Parameters:
  • parser (ArgumentParser)

  • tool (str)

  • tracks (str)

  • variant_help (str)

Return type:

None

command_text(argv=None)#

The invocation, as a copy-pasteable line, with credentials masked.

argv defaults to this process’s own sys.argv. Pass another process’s argv when the run is opened somewhere the user never typed a command – a worker subprocess, whose own sys.argv is an implementation detail rather than a reproducible invocation.

Parameters:

argv (list[str] | None)

Return type:

str

current_user()#

Return the current username, or "unknown" if the uid has no passwd entry.

Return type:

str

default_experiment_name(tool, model, variant, user=None)#

Build an experiment name of the form <user>/<tool>/<model>-<variant>.

Only the basename of model is used, so a local checkpoint directory and an org/name Hugging Face id collapse to the same readable name; variant is whatever distinguishes this run of tool on model, such as a recipe name or a quantization format. Each component is reduced to [A-Za-z0-9._-] so the / separators stay meaningful, and user defaults to the current user.

Example

>>> default_experiment_name("hf_ptq", "/models/Qwen3-0.6B/", "nvfp4", user="alice")
'alice/hf_ptq/Qwen3-0.6B-nvfp4'
Parameters:
  • tool (str)

  • model (str)

  • variant (str)

  • user (str | None)

Return type:

str

drop_experiment_json(checkpoint_dir)#

Remove a provenance pointer an untracked export would otherwise inherit.

A fresh checkpoint written into a reused output directory would keep the previous run’s pointer, and one produced from a tracked source checkpoint could be handed that source’s pointer. Either way the file would name a run that did not produce these weights. Call it only for a completed export; a failed run leaves whatever checkpoint was already there, pointer included.

Parameters:

checkpoint_dir (Path | str)

Return type:

None

resolve_mlflow_args(args, parser, tool, model, variant)#

Settle where tracking is configured from, and name the experiment, in place.

Sets args.mlflow to the validated URI or None, args.mlflow_required to whether the flag asked for it, and defaults args.mlflow_experiment from tool, model and variant. Pair with add_mlflow_args().

Parameters:
  • args (Namespace)

  • parser (ArgumentParser)

  • tool (str)

  • model (str)

  • variant (str)

Return type:

None

resolve_tracking_uri(uri, parser)#

Settle the tracking URI from --mlflow and the environment.

Returns (uri or None, required), where required records that the flag was passed. Only the flag is a deliberate request, so only the flag is fatal when the URI is unusable: the environment variable is commonly exported for unrelated tooling and must not fail a job that would otherwise have worked.

Parameters:
  • uri (str | None)

  • parser (ArgumentParser)

Return type:

tuple[str | None, bool]

validate_tracking_uri(uri)#

Validate an MLflow tracking URI and return it without a trailing slash.

Only http(s) servers are accepted; MLflow’s local file: / sqlite: backends are not a useful destination for a shared record of a run.

Raises:

ValueError – If uri is empty, has no host, or is not an http(s) URL.

Parameters:

uri (str)

Return type:

str