Skip to content

Artifact storage

Updated 3 min read

MLflow stores two different things in two different places:

Backend storeArtifact store
Holdsruns, params, metrics, tags, the model registrymodel files, plots, datasets, anything log_artifact writes
This chart’s defaultbundled PostgreSQL with an 8Gi PVC./mlruns, a path inside the MLflow container
Survives a pod restartyesno

The backend store is configured for durability. The artifact store is not.

Configure a bucket before people start logging models.

mlflow:
artifactRoot:
proxiedArtifactStorage: true
s3:
enabled: true
bucket: my-mlflow-artifacts
path: "" # optional prefix
existingSecret:
name: mlflow-s3-credentials
keyOfAccessKeyId: AWS_ACCESS_KEY_ID
keyOfSecretAccessKey: AWS_SECRET_ACCESS_KEY

Create the credentials secret separately, in the release namespace:

Terminal window
kubectl -n mlflow create secret generic mlflow-s3-credentials \
--from-literal=AWS_ACCESS_KEY_ID=... \
--from-literal=AWS_SECRET_ACCESS_KEY=...

The upstream chart also accepts awsAccessKeyId and awsSecretAccessKey inline. Use the existing-secret form on any cluster whose values live in git.

For IRSA or another pod-identity mechanism, omit the credentials entirely and annotate the service account with the role ARN.

S3-compatible storage (MinIO, Hetzner, Backblaze)

Section titled “S3-compatible storage (MinIO, Hetzner, Backblaze)”

Same s3 block, plus an endpoint override:

mlflow:
extraEnvVars:
MLFLOW_S3_ENDPOINT_URL: http://minio.minio.svc.cluster.local:9000
AWS_DEFAULT_REGION: us-east-1
mlflow:
artifactRoot:
proxiedArtifactStorage: true
gcs:
enabled: true
bucket: my-mlflow-artifacts
path: ""
mlflow:
artifactRoot:
proxiedArtifactStorage: true
azureBlob:
enabled: true
container: mlflow-artifacts
storageAccount: mystorageaccount
connectionString: "" # or accessKey

proxiedArtifactStorage: true routes artifact uploads and downloads through the MLflow server, which holds the bucket credentials. Clients need no bucket access of their own.

That is the right default here. JupyterHub notebooks already reach the tracking server over the cluster network; giving every notebook pod its own bucket credentials would be a larger blast radius for no gain.

The cost is that artifact traffic passes through the MLflow pod, so large model uploads consume its bandwidth and memory. For very large artifacts with clients that can be trusted with credentials, set it to false and let clients talk to the bucket directly.

After the change, log something and confirm it landed:

import mlflow
mlflow.set_experiment("artifact-check")
with mlflow.start_run() as run:
with open("hello.txt", "w") as f:
f.write("hi")
mlflow.log_artifact("hello.txt")
print(mlflow.get_artifact_uri())

With proxiedArtifactStorage: true the printed URI is mlflow-artifacts:/<experiment>/<run>/artifacts — the client talks to the tracking server, which is the one holding the bucket credentials. That is the success case; a plain local filesystem path is the failure case. Confirm the object actually landed by looking in the bucket. With proxying off, the URI is the bucket scheme directly (s3://, gs://, wasbs://).

The server’s artifact configuration is on its command line, not in its environment — the chart passes --artifacts-destination (proxied) or --default-artifact-root (direct) as flags:

Terminal window
kubectl -n mlflow get deploy mlflow-pack \
-o jsonpath='{.spec.template.spec.containers[0].args}'

There is no migration path from the ephemeral default, because by the time you notice, the files are usually already gone. If a pod is still running with artifacts you care about, copy them out before changing anything:

Terminal window
kubectl cp mlflow/<pod>:/mlflow/mlruns ./mlruns-backup

Then reconfigure the artifact root. Existing runs keep their recorded artifact URIs, so they will still point at the old local paths — new runs use the bucket. Re-logging is usually simpler than rewriting URIs in the database.

The upstream chart notes that autoscaling is only supported when the backend store is not SQLite and the artifact root is one of the blob storage backends. With the default local artifact root, replicas would each hold a different subset of artifacts. Configuring a bucket is a prerequisite for running more than one MLflow replica.