Contents
Figure 1: One compose.yaml instead of six terminal windows, each running a different piece of your ML stack
Your ML project isn't just a training script anymore. It's a model server, a vector database, a feature store, maybe a Jupyter environment, and increasingly, a GPU that needs to be shared intelligently across some of those pieces. Docker Compose is how you wire all of that together into something you can start with one command and actually reproduce on a teammate's machine.
I'll walk through the current 2026 patterns specifically, since GPU support, watch mode, and profiles have genuinely changed how Compose gets used for ML work compared to a couple years ago.
If you've been following this series' MLOps for beginners guide, think of this as the local-development half of that lifecycle — the environment where your CI/CD pipeline's builds and tests actually get reproduced before anything ships.
A Quick Naming Note
If you see docker-compose (with a hyphen) in an old tutorial, know that the standalone binary is deprecated. Compose functionality is now integrated directly into the docker compose command (no hyphen, as a subcommand of the Docker CLI). Install it alongside Docker with sudo apt-get install docker-compose-plugin if it's not already there, and use docker compose up, not docker-compose up, going forward.
Why Compose Fits ML Projects So Well
A typical ML development stack has genuinely different services with different resource needs: a GPU-hungry inference server, a lightweight API gateway, a database that needs persistent storage, maybe a monitoring dashboard. Compose lets you define all of them declaratively in one file, start them together with correct startup ordering, and tear the whole thing down cleanly when you're done.
For local development, CI pipelines, staging environments, and single-machine GPU workloads, it's become genuinely hard to argue against Compose as the default choice, even as Kubernetes dominates the production conversation for larger deployments.
GPU Support: The Part Everyone Gets Wrong First
Before any Compose file works with a GPU, your host machine needs the right NVIDIA drivers and the NVIDIA Container Toolkit installed, not just Docker itself. Run nvidia-smi on the host first — if that doesn't show your GPU, nothing downstream will work either, no matter how correct your Compose file is.
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
Verify it actually works before writing a single line of Compose:
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
If that command prints your GPU from inside a container, the host side is done — everything after this point is just YAML.
Declaring GPU Access in Your Compose File
GPUs are referenced using the device attribute from the Compose Deploy specification:
services:
ml-training:
image: my-pytorch:cuda12.1
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
volumes:
- ./data:/workspace/data
- ./models:/workspace/models
environment:
- CUDA_VISIBLE_DEVICES=0
command: python3 train.py
capabilities: [gpu] is required — that field must be set. count: all grabs every available GPU; use a specific number (count: 1) to limit how many a single service claims.
Targeting Specific GPUs on a Multi-GPU Host
On a machine with multiple GPUs, device_ids lets you pin specific services to specific cards, which matters a lot once you're running more than one GPU-bound service side by side:
services:
embedding-service:
image: my-embedding-model:latest
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ['0']
capabilities: [gpu]
inference-server:
image: my-llm-server:latest
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ['1', '2']
capabilities: [gpu]
| Option | Use It When | Behavior |
|---|---|---|
count: all | One service should own the whole machine | Claims every visible GPU |
count: N | You want a hard cap per service | Claims the next N available GPUs |
device_ids: ['0'] | Services must map to specific cards | Pins the service to exact GPUs |
This explicit device reservation is what prevents conflicts between services when you've got a genuinely mixed workload — an embedding model on GPU 0, a larger LLM spanning GPUs 1 and 2, without either one silently stepping on the other's memory.
A Realistic Multi-Container ML Stack
Here's how these pieces come together for a local RAG-style development setup:
services:
ollama:
image: ollama/ollama:latest
volumes:
- ollama_data:/root/.ollama
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
ports:
- "11434:11434"
vector-db:
image: qdrant/qdrant:latest
volumes:
- qdrant_data:/qdrant/storage
ports:
- "6333:6333"
api:
build: ./api
depends_on:
ollama:
condition: service_healthy
vector-db:
condition: service_started
environment:
- OLLAMA_HOST=http://ollama:11434
- QDRANT_HOST=http://vector-db:6333
ports:
- "8000:8000"
volumes:
ollama_data:
qdrant_data:
Note the service names (ollama, vector-db) doubling as hostnames other containers use to reach them — that's Compose's built-in networking, and it's a big part of why this pattern works so cleanly for multi-service local development.
If you've never run this stack in pieces before, the Ollama setup guide covers the model server itself, and the Qdrant tutorial covers the vector database — this article is the glue that runs them together. The finished shape looks a lot like the local RAG chatbot we built earlier in this series, just containerized.
Health Checks and Startup Ordering
depends_on alone only controls start order, not readiness. A model server container can report "started" long before it's actually finished loading weights and ready to accept requests. Pair it with a proper health check:
services:
ollama:
image: ollama/ollama:latest
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:11434/api/tags"]
interval: 10s
timeout: 5s
retries: 5
start_period: 30s
api:
build: ./api
depends_on:
ollama:
condition: service_healthy
The condition: service_healthy clause is what actually makes api wait for Ollama to be genuinely ready, not just running. This single pattern eliminates a huge class of "works on retry, fails on first start" bugs that plague ML dev environments specifically, since model loading times are rarely instant.
Watch Mode: Killing the Rebuild Cycle
If you're iterating on application code around your model, not the model itself, rebuilding a container image on every change gets old fast. Compose's watch mode syncs code changes into a running container without a full rebuild:
services:
api:
build: ./api
develop:
watch:
- action: sync
path: ./api/src
target: /app/src
- action: rebuild
path: ./api/requirements.txt
Run it with docker compose watch. Code changes under ./api/src sync instantly into the running container; changes to requirements.txt trigger a full rebuild, since new dependencies genuinely need a fresh image. This distinction — sync for code, rebuild for dependencies — is what actually made container-based ML development feel fast rather than sluggish.
Profiles: One File, Multiple Environments
Not every service belongs in every environment. Profiles let you tag services and selectively activate them:
services:
api:
build: ./api
# always runs, no profile needed
gpu-training:
image: my-pytorch:cuda12.1
profiles: ["training"]
deploy:
resources:
reservations:
devices:
- driver: nvidia
capabilities: [gpu]
mock-payment-api:
image: my-mock-payments:latest
profiles: ["dev"]
docker compose --profile training up # includes gpu-training
docker compose up # skips both profiled services
This is genuinely useful for keeping a mock payment API or a heavy GPU training service out of your default docker compose up, so a teammate without a GPU, or without a reason to hit a mock payment endpoint, doesn't accidentally pull in services they don't need. One team reported this exact pattern saved them from accidentally hitting a live payment API during development more than once — keeping risky or heavy services behind an explicit profile flag is a genuinely cheap safeguard.
Resource Limits Matter for ML Workloads Specifically
GPU reservations get the attention, but CPU and memory limits matter too, especially when several services share one development machine:
services:
my-service:
deploy:
resources:
limits:
cpus: '2.0'
memory: 4G
reservations:
cpus: '1.0'
memory: 2G
limits caps what a service can consume; reservations guarantees a minimum. Without these, one runaway preprocessing job can starve your model server of memory on a shared dev box — a frustrating failure mode that's genuinely easy to prevent with a few lines of YAML.
Bridging to Production With Bake
Compose is primarily a development and single-machine tool, but Docker Bake helps bridge the gap between your local Compose workflow and building production-ready images, letting you define build configurations that stay consistent between your docker compose build locally and your CI/CD image builds. If your team's production deployment target is Kubernetes rather than Compose itself, Bake keeps the image-building step consistent across both without duplicating build logic in two places.
Managed pipelines like the SageMaker Pipelines setup pick up roughly where Compose leaves off — same containers, executed as tracked, auditable steps on AWS infrastructure instead of your laptop.
A Practical Decision Framework
- Verify the host first:
nvidia-smion the machine, then the NVIDIA Container Toolkit — no Compose YAML fixes a missing driver - Write one
compose.yamlwith your model server, vector database, and API as separate services - Add health checks to anything that loads a model, and use
condition: service_healthyon its dependents - Reserve GPUs explicitly with
capabilities: [gpu], usingcountfor simple cases anddevice_idsto pin multi-service workloads - Set CPU and memory limits on every service that shares a dev machine
- Put heavy or risky services behind profiles so plain
docker compose upstays safe and light - Turn on watch mode when you're iterating on application code, keeping rebuilds for dependency changes only
- Keep builds consistent with Bake when CI and production need the same image your laptop just built
The ordering matters more than the individual steps — host verification and health checks prevent the two failure modes that make teams abandon Compose for ML work in the first place.
Common Mistakes People Make
Forgetting host-level GPU setup
Recall the GPU section — Compose file GPU declarations do nothing if nvidia-container-toolkit isn't installed and configured on the host first. Test with the one-line docker run --gpus all check before debugging YAML.
Relying on depends_on alone for readiness
Add proper health checks, especially for anything that loads a model on startup. Start order is not readiness, and ML services have unusually long startup times compared to typical web services.
Skipping resource limits on shared dev machines
One heavy service without limits can starve everything else running alongside it. A few lines of YAML prevents the most annoying shared-box failure mode there is.
Using the deprecated docker-compose binary
FYI: switch to docker compose (the plugin) if you haven't already. Old tutorials will keep steering you wrong for a while; the hyphenated binary isn't getting new features.
Not using profiles for optional or risky services
Heavy GPU services and anything that could hit a real external API belong behind an explicit profile flag — it's the cheapest safety rail in this entire article.
Recommended Books
- Docker Deep Dive by Nigel Poulton — the clearest starting point for how containers and Compose actually work under the images, volumes, and networking concepts this article leans on.
- Docker for Data Science by Joshua T. Perkins — specifically about containerizing data and ML workflows, which is exactly the gap between generic Docker tutorials and the stack above.
- Docker: Up & Running by Karl Matthias and Brendan Smith — the production-minded companion, useful when your Compose setup graduates toward the Bake/CI workflow at the end of this article.
Want to Go Deeper?
If you want structured practice on containerized ML workflows, Educative's ML courses include hands-on labs that pair well with this kind of environment work. The unlimited plan is useful when you're working through several infrastructure topics in one stretch.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is Docker Compose used for in ML projects?
Compose defines a multi-service ML stack — model server, vector database, API, monitoring — in one YAML file, starts them together with correct ordering and health-aware readiness, and tears the whole environment down cleanly. It is the standard for local development, CI pipelines, staging, and single-machine GPU workloads.
How do you enable GPU access in a Docker Compose file?
Declare devices under each service's deploy.resources.reservations.devices with driver: nvidia, capabilities: [gpu], and either count: all or a specific count. The host must also have NVIDIA drivers and the nvidia-container-toolkit installed and configured for Docker first — the Compose declaration alone does nothing without it.
What is the difference between docker-compose and docker compose?
docker-compose (with a hyphen) is the deprecated standalone binary. Compose is now integrated into the Docker CLI as a subcommand: docker compose (no hyphen), installed via the docker-compose-plugin package. New projects should use docker compose up.
Why isn't depends_on enough for model server containers?
depends_on only controls start order, not readiness. A model server can report started long before it has finished loading weights. Pair it with a healthcheck and condition: service_healthy so dependent services wait for genuine readiness instead of racing a cold model load.
What does Docker Compose watch mode do?
docker compose watch syncs code changes into a running container without rebuilding the image, and can be configured to trigger a full rebuild when dependency files like requirements.txt change. It removes the rebuild-on-every-edit cycle when iterating on application code around a model.
When should you use Compose profiles?
Profiles tag services so they only start when explicitly requested with --profile. Use them for heavy GPU training services, mock external APIs, or anything a teammate without the right hardware or permission shouldn't pull in with a plain docker compose up.
Wrapping This Up
Docker Compose in 2026 handles real ML development workloads well: GPU device reservations with per-service targeting, watch mode for fast iteration, profiles for managing multiple environments from one file, and health checks that prevent the classic "started but not ready" race condition.
Will Compose replace Kubernetes for your production ML platform? No, and it's not trying to. But for local development, CI, staging, and single-machine GPU workloads, spinning up your model server, vector database, and API together with one command — and tearing it all down just as cleanly — is hard to beat. Wire up your next ML project's services in one compose.yaml tonight, and you'll never go back to juggling separate terminal windows for each piece.
Once it's running, the pieces you've containerized slot straight into the rest of this series' stack — benchmark the model server inside its container, then hand the same image to CI/CD for automated testing.
Related Articles
- Run LLMs Locally with Ollama: Complete Setup Guide (2026)
- Building a Local RAG Chatbot with Ollama and LangChain
- Qdrant Tutorial: Open-Source Vector Search Engine Getting Started
- CI/CD for Machine Learning: Automate Model Testing and Deployment (2026)
- MLOps for Beginners: Complete Guide to Production Machine Learning