Contents
Figure 4: DVC keeps your Git repository lean while your actual data lives safely in cloud storage — versioned, traceable, reproducible
The MLOps beginners guide named the single most-skipped stage in the whole lifecycle: "you can't reproduce a model's behavior if you can't reproduce the exact data it trained on." Git solves this beautifully for code. Git chokes badly on a 40GB dataset or a 2GB model checkpoint — try committing either and watch your repository slow to a crawl or outright refuse the push. DVC exists specifically to fix that gap, and it does it by not actually storing your data in Git at all.
This closes the versioning piece of the data-and-MLOps infrastructure arc running through the last several articles — extraction (the ETL tutorial), transformation (dbt), orchestration (Airflow), feature consistency (Feast), and now, the piece making all of it reproducible after the fact: knowing with certainty which exact version of your data produced which exact model.
By the end of this guide, you'll understand DVC's actual mechanism — how it tracks huge files without bloating Git — build a versioned dataset, and construct a reproducible pipeline using its dvc.yaml format. IMO, the moment dvc add clicks conceptually is genuinely one of the more satisfying "oh, that's clever" realizations in this whole series :)
The Core Mechanism: Pointers in Git, Payload Elsewhere
DVC is a command-line tool and Apache 2.0 open-source library for versioning datasets, models, pipelines, and ML experiments on top of Git — genuinely "Git for data and models," built to complete the reproducibility story that Docker and virtual environments only partially solve.
Docker and virtual environments solve two of three reproducibility problems — consistent runtime environment, consistent dependencies. DVC completes the third: consistent data and model artifacts across time and across team members.
DVC is not a network service — there's no DVC server, no REST API. It stores small pointer files in Git, while the actual heavy payload (your dataset, your model checkpoint) lives in a remote you already own: S3, GCS, Azure, SSH, HDFS, or even a local drive.
This means DVC's own cost is genuinely just your storage backend — there's no DVC subscription fee, no vendor lock-in beyond whichever cloud storage you were probably already paying for.
How dvc add Actually Works
pip install dvc
git init
dvc init
dvc add data/raw/imdb_sample.csv
This single command does something genuinely clever: it moves your actual CSV into DVC's local cache, replaces it in your working directory with the same file (via a link), and creates a tiny .dvc metadata file — imdb_sample.csv.dvc — containing a content hash and the file's size. That small metadata file is what actually gets committed to Git, not the dataset itself.
git add data/raw/imdb_sample.csv.dvc data/raw/.gitignore
git commit -m "Track raw IMDb sample dataset"
Notice DVC automatically adds the real data file to .gitignore — Git genuinely never sees the actual dataset, only the pointer describing it. This is the entire mechanism that keeps your repository small and fast regardless of how large your underlying data grows.
Pushing Data to Remote Storage
The .dvc file alone doesn't back up your actual data — for that, you need a configured remote, exactly the same concept as a Git remote, just for the data payload instead of code.
dvc remote add -d myremote s3://my-bucket/dvc-storage
dvc push
dvc push uploads the actual cached data to that remote — S3 here, but Azure, GCS, SSH, or a local network drive all work identically. Teammates cloning your Git repo get the lightweight .dvc pointer files immediately, then run dvc pull to fetch the actual multi-gigabyte payload only when they genuinely need it.
Switching Between Data Versions Like Git Branches
This is genuinely where the "version control" framing earns its name rather than being a metaphor.
git checkout v1.0-dataset
dvc checkout
git checkout switches which .dvc pointer file is active; dvc checkout then syncs your actual working files to match that pointer — swapping in the exact dataset version that existed at that commit. This is genuinely the mechanism that lets you answer "which data produced this specific model" with total confidence, months later, rather than guessing based on a folder name or a Slack message nobody can find anymore.
DVC Pipelines: Making Reproduction Automatic, Not Manual
Beyond versioning individual files, DVC lets you define entire multi-step pipelines — preprocessing, training, evaluation — as a DAG, genuinely the same dependency-graph concept from the Airflow and dbt articles, expressed in DVC's own YAML format instead.
# dvc.yaml
stages:
preprocess:
cmd: python preprocess.py
deps:
- data/raw/imdb_sample.csv
- preprocess.py
outs:
- data/processed/train.csv
train:
cmd: python train.py
deps:
- data/processed/train.csv
- train.py
params:
- learning_rate
- epochs
outs:
- models/model.pkl
metrics:
- metrics.json:
cache: false
dvc repro
dvc repro is genuinely the standout feature here: it only reruns the stages actually impacted by whatever changed — if you only edited train.py, the preprocess stage's cached output gets reused untouched, and only train reruns. This directly mirrors the "idempotent, retry-only-what-failed" principle from the Airflow article, just applied to local pipeline reproducibility rather than production orchestration.
Experiment Tracking Without a Server
Recall the MLOps beginners guide introducing MLflow for experiment tracking. DVC offers a genuinely different, server-less angle on the same underlying problem — turning your own machine into an experiment management platform using nothing but Git underneath.
dvc exp run --set-param learning_rate=0.001
dvc exp run --set-param learning_rate=0.01
dvc exp show
dvc exp show displays a comparison table across every experiment run — parameters, metrics, and which code/data version produced each — genuinely comparable to MLflow's tracking UI, but stored entirely as Git commits rather than requiring a separate tracking server or database. Choose DVC's experiment tracking when you want zero additional infrastructure; choose MLflow when you want a richer, dedicated UI — the two aren't mutually exclusive, and plenty of teams genuinely use both for different purposes.
Codification: DVC's Actual Design Philosophy
Worth stating explicitly, since it's genuinely the philosophy tying every feature above together: DVC's core design principle is codification — defining every aspect of an ML project (data versions, model versions, pipelines, experiments) in human-readable metafiles, so that established software engineering practices and toolsets apply directly to data science work, rather than data science living in a separate, less rigorous world of notebooks and shared spreadsheets.
DVC explicitly aims to replace the spreadsheet-and-shared-document approach many teams default to for tracking "which dataset version produced which model" — the kind of ad-hoc tracking that works fine for one person and collapses the moment a second person joins the project.
DVC Does Not Replace Git — It Extends It
Worth being precise about this, since it's a genuinely common point of confusion: DVC is not fundamentally bound to Git and can technically work without it (except for its versioning-related features specifically), but in essentially every real deployment, DVC builds directly on top of Git rather than replacing it.
Git still tracks your code, your dvc.yaml pipeline definitions, and the small .dvc metadata pointer files.
DVC handles exclusively the large binary payload — datasets, model checkpoints — that Git was never designed to store efficiently.
This division of labor is genuinely deliberate, not a limitation — each tool does the one job it's actually built for, echoing the same "specialized tools, not one tool that does everything" principle from the MLOps platforms comparison article.
CI/CD Integration: CML
DVC integrates with CML (Continuous Machine Learning) specifically for training models in the cloud or on Kubernetes as part of a CI/CD pipeline — genuinely relevant if you've followed the Airflow article's version-control-your-DAGs advice and want the same discipline extended to automated retraining triggered by a pull request or a new data version landing.
Where DVC Fits in This Series' Full Data-and-MLOps Stack
Recall the complete pipeline built across recent articles: extraction/loading, dbt transformation, Airflow orchestration, Feast for training/serving feature consistency. DVC is genuinely the piece answering a different question than any of those — not "how does data flow" but "which exact version of this data and model combination am I looking at, and can I get back to it later."
Raw data → ETL/dbt (transform) → DVC (version the training dataset)
↓
dvc repro (train, versioned)
↓
Model + metrics, all traceable to exact data version
A dataset produced by your dbt gold-layer models is genuinely a great candidate for dvc add the moment it's exported for training — giving you a permanent, Git-trackable record of exactly which transformed dataset snapshot a given model checkpoint was trained against.
Common Mistakes People Make
Forgetting to commit the .dvc and dvc.yaml files to Git. These are the actual pointers making versioning work — skip committing them, and you've lost the entire reproducibility chain despite having run every DVC command correctly.
Never configuring or pushing to a remote. Data tracked only in your local DVC cache isn't backed up or shareable — dvc push to a real remote is genuinely not optional for team collaboration.
Trying to commit large files directly to Git instead of using dvc add. This is precisely the problem DVC exists to prevent — a multi-gigabyte file committed straight to Git will bloat and slow your repository in exactly the way the tool was built to avoid.
Skipping DVC commands in CI/CD pipelines. Automating data and model handling requires actually integrating dvc pull/dvc repro into your pipeline, not just running them manually on your own machine.
Treating DVC as a replacement for Git rather than an extension of it. Git still owns your code and pipeline definitions — DVC's job is specifically the large binary artifacts Git handles poorly.
Recommended Books
- Designing Machine Learning Systems by Chip Huyen — the definitive guide to production ML, including data versioning's critical role in reproducibility and exactly the gap DVC fills.
- Practical MLOps by Noah Gomez — hands-on coverage of DVC, MLflow, and the reproducibility tooling that makes ML systems reliable in production rather than experimental notebooks.
- Machine Learning Engineering by Andriy Burkov — rigorous treatment of ML pipeline design and reproducibility principles that DVC's architecture directly implements.
- Data Versioning with DVC — community reference covering advanced DVC patterns, pipeline design, and integration with the broader MLOps stack.
Wrapping This Up
DVC solves the exact reproducibility gap the MLOps beginners guide flagged as most commonly skipped: it stores lightweight, content-hashed pointer files in Git while the actual large datasets and model checkpoints live in a remote you already own, giving you Git's full versioning discipline — branches, commits, checkouts — applied to artifacts Git itself was never built to handle efficiently. dvc repro's selective re-execution and dvc exp experiment tracking extend that same discipline to entire pipelines and experiment comparisons, all without requiring a dedicated tracking server.
Remember that DVC doesn't replace Git — it extends it specifically for the large binary artifacts Git handles poorly — and that forgetting to push to a configured remote leaves your versioned data unshared and unbacked-up despite everything else working correctly. FYI, this genuinely completes the reproducibility side of the data-and-MLOps arc running through this series' recent articles: you can now trace a model back through its exact training data version, the dbt transformation that produced it, and the pipeline stage that generated it — the full chain the very first MLOps article argued was the actual difference between a model that ships and one that stays stuck in a notebook forever :)
Now go take the daily_sales_summary output from the ETL pipeline article, run dvc add on it, and commit the resulting .dvc file to Git. That's genuinely the smallest possible step that turns "a CSV sitting in a folder" into "a dataset with a permanent, recoverable version history."