MLflow

MLflow

Open Source

Deliver high-quality AI, fast.

Data & ML Libraries
ML Operations

Published 29 May 2026 · Last updated 27 September 2026

Scores

Popularity4/5

The most widely adopted ML experiment tracking tool — 20K+ GitHub stars, 60M+ monthly PyPI downloads, deeply embedded in the MLOps ecosystem. Dominant within the experiment tracking niche but that niche is narrow; minimal presence outside the ML engineering community.

Learning Curve3/5

Basic experiment tracking (autolog + UI) takes minutes to learn. Full setup with model registry, multi-user auth, and deployment integrations has a meaningful ramp that goes beyond the initial quick start.

Flexibility4/5

Highly pluggable: swappable tracking stores (SQLite, PostgreSQL, MySQL), artifact backends (S3, GCS, Azure Blob, SFTP), and custom evaluation judges via plugin API. Very configurable, though conventions around the tracking server and run lifecycle still apply.

Performance4/5

Handles hundreds of concurrent runs and multi-GB artifact storage reliably. Gateway server was merged into the tracking server in v3.9 to reduce overhead. Minor UI slowdowns at very high run counts.

Portability5/5

Completely self-hostable, cloud-agnostic, backend-agnostic, and open-source licensed. Managed option available via Databricks but no lock-in — teams can migrate tracking stores and artifact backends without touching their training code.

About MLflow

MLflow is an open-source platform for the machine learning lifecycle, created at Databricks in 2018 and now a Linux Foundation project under the Apache 2.0 license. It is framework-agnostic and made of components that work on their own or together.

Tracking logs parameters, metrics, code versions, and artifacts for every training run through a small Python API, and mlflow.autolog() captures them automatically for PyTorch, scikit-learn, XGBoost, TensorFlow, and other libraries. A web UI compares runs side by side and charts metrics over time. The Model Registry stores versioned models with lineage back to the run that produced them, and aliases (such as "champion") and tags mark which version a deployment should load. MLflow Models is a standard packaging format that serves a model as a REST endpoint or deploys it to SageMaker, Azure ML, Spark batch jobs, or a Docker container.

MLflow 3 extended the platform to generative AI. Tracing records every step of an LLM or agent call (prompts, tool calls, retrieved documents, latency, and cost) with OpenTelemetry-compatible spans. Evaluation runs LLM judges and custom scorers over datasets, a prompt registry versions prompts, and the AI Gateway puts several LLM providers behind one API with credentials and rate limits managed centrally.

The tracking server self-hosts with a single command, starting on SQLite and local files and scaling to PostgreSQL or MySQL with S3, GCS, or Azure Blob for artifacts. Managed MLflow is offered by Databricks, Amazon SageMaker, and other platforms for teams that would rather not run it.

Key Features

  • Experiment tracking with autologging for major ML frameworks
  • Model Registry with versions, aliases, tags, and lineage
  • Standard model packaging for REST serving, SageMaker, Azure ML, and Spark
  • Tracing for LLM and agent applications, OpenTelemetry-compatible
  • LLM evaluation with built-in judges and custom scorers
  • Prompt registry for versioning prompts
  • AI Gateway for a single API across LLM providers
  • Self-hostable with SQL database and object storage backends

Pros

  • Adds tracking to existing code in a few lines with autologging
  • Works with any ML framework and any cloud
  • Covers classic ML and GenAI tracing and evaluation in one tool
  • Scales from a local SQLite setup to a shared team server
  • Managed versions on Databricks and SageMaker avoid running it yourself

Cons

  • A shared team server needs auth, storage, and TLS set up by hand
  • No pipeline orchestration; scheduling needs Airflow, Prefect, or Dagster
  • UI slows down when comparing hundreds of runs with many metrics
  • Authentication is off by default on a self-hosted server

MLflow Pricing

Open Source

Tech Stacks with MLflow

MLOps Pipeline

Project

Production-grade ML infrastructure. PyTorch for model training, Apache Airflow (or Dagster or Prefect) for orchestration, dbt for feature transformations, and Snowflake as the data warehouse, with Docker as an optional containerization addition.

Deploy on:
Orchestrator:
Data Libraries:
Model Serving (Python API):
Experiment Tracking add-on:
CI/CD add-on:
Containerization add-on:

PyTorch ML Training

Project

Train deep learning models with PyTorch, with scikit-learn baselines to compare against, Pandas for data preparation, and Jupyter for experimentation. MLflow or Weights & Biases can be added to track experiments and model versions once runs need comparing.

Experiment Tracking add-on:
CI/CD add-on:
Containerization add-on:

TensorFlow ML Training

Project

Build and train ML models using TensorFlow and Keras, from prototyping in Jupyter notebooks to exported models ready for serving. MLflow or Weights & Biases can be added to track experiments and model versions once runs need comparing.

Experiment Tracking add-on:
CI/CD add-on:
Containerization add-on:

Tools Related to MLflow

Works well with MLflow(3)

A common production pattern loads an MLflow model from the registry at startup and serves it behind a FastAPI endpoint — this separates model versioning (MLflow registry) from API design and request validation (FastAPI).

Apache Airflow and MLflow are a common MLOps pairing: Airflow DAGs schedule and orchestrate training jobs, while MLflow tracks the experiment results, model versions, and artifacts produced by each DAG run.

Prefect and MLflow are commonly paired: Prefect flows orchestrate ML training pipelines (scheduling, retries, parameter injection) while MLflow tracks the experiments, metrics, and model artifacts produced by each run.

Integrates with MLflow(7)

MLflow has official integration with AWS — S3 is a first-class artifact store backend, and mlflow.sagemaker.deploy() can push a registered model directly to an AWS SageMaker real-time inference endpoint.

Databricks created MLflow and offers Managed MLflow — Databricks notebooks, Jobs, and Unity Catalog integrate natively with MLflow tracking and the model registry, making it the fully managed option for MLflow at enterprise scale.

Dagster provides a first-party dagster-mlflow integration library — Dagster assets and ops can automatically log parameters and metrics to MLflow, and the MLflow model registry can be used as a Dagster artifact store.

MLflow provides official autologging for scikit-learn — mlflow.sklearn.autolog() captures estimator parameters, cross-validation metrics, and the trained model with one line of code; the most common MLflow entry point for classical ML.

MLflow provides first-party autologging for PyTorch — mlflow.pytorch.autolog() captures training metrics, hyperparameters, and model checkpoints automatically, and mlflow.pytorch.log_model() packages the model for registry and serving.

MLflow works with TensorFlow — mlflow.tensorflow.autolog() captures metrics, parameters, and SavedModel artifacts automatically during TF training runs.

Alternatives to MLflow(1)

MLflow and Weights & Biases are the two dominant ML experiment tracking platforms — MLflow is preferred for self-hosted setups and production deployment integration, while W&B is preferred for ease of setup and rich collaborative dashboards.

Tags

PythonOpen SourceSelf-hostableDocker CompatibleMachine LearningData EngineeringData PipelinesData ScienceWeb

Details

Maintained
Yes