Why a checklist before production
Docker Compose was designed for development: its defaults favor simplicity over robustness. A container will not restart after a host reboot, its logs grow without limit, it can consume all the machine's RAM, and it listens on every network interface. In development, none of these behaviors cause problems because you relaunch the stack by hand several times a day. In production, these missing settings turn into incidents: full disk at 3 a.m., database lost after a docker compose down, service unreachable after a power outage. The good news: hardening a Compose stack requires no rewrite, only a dozen targeted additions. Run through each point before going live, and your service will handle load and reboots without constant monitoring.
The 10 points at a glance
- restart: unless-stopped — the container restarts after a host reboot or crash
- healthcheck — Docker detects a stuck container and enables rolling restarts
- CPU/memory limits — a runaway service can no longer starve its neighbors
- secrets via .env or Docker secrets — never a plaintext password in the Compose file
- named volumes — data survives a
docker compose down - log rotation —
max-sizeandmax-fileprevent disk saturation - network isolation — separate
frontendandbackend, don't expose everything on the default bridge - port binding —
127.0.0.1:PORTbehind a reverse proxy, not0.0.0.0 - pinned image tags — a version or digest, never
latest - depends_on with condition —
service_healthyprevents startup races
Prerequisites
Before applying this checklist, make sure you have a VPS with Docker Engine and the Compose v2 plugin installed (the command is docker compose, without a hyphen, since 2022). Check the version with docker compose version: the deploy.resources syntax outside Swarm requires Compose v2. Place your file in a dedicated project folder, for example /opt/myapp, with a .env file alongside it and restricted permissions (chmod 600 .env). Plan for an upstream reverse proxy — Traefik, Caddy or Nginx — since several settings, especially port binding, assume that public traffic never touches your containers directly. Finally, keep a copy of your Compose file under version control.
The 10 steps in detail
1. Restart policy
Add
restart: unless-stoppedto every service. The container restarts after a crash or host reboot, but stays stopped if you intentionally stopped it. Avoidrestart: alwayswhich would relaunch even a container you wanted to leave stopped, and preferunless-stoppedfor application services andon-failurefor one-off tasks that should not loop indefinitely.2. Healthcheck
Declare a
healthcheck:block with atest(e.g.curl -f http://localhost:8080/health || exit 1), aninterval(check frequency, e.g.30s), atimeout(failure threshold, e.g.10s) andretries(consecutive failures before switching tounhealthy, e.g.3). Also addstart_period: 40sfor services that take time to boot, to avoid marking themunhealthyduring initialization. Docker then marks the containerhealthyorunhealthy, allowing other services to react to a stall viadepends_on: condition: service_healthy.3. Resource limits
Under
deploy.resources.limits, setcpusandmemory(e.g.memory: 512M). Without a limit, a leaking service can consume all RAM and cause the OOM killer to terminate others. Also addreservationsto guarantee a minimum at startup. In Compose v2 outside Swarm,limitsare applied as cgroup constraints;reservationsare declarative and do not block startup, but serve as a planning reference.4. Secrets outside the file
Never put a plaintext password in the YAML. Reference them via
env_file: .envor${VARIABLE}, or use Docker'ssecrets:mechanism which mounts the secret as a file in/run/secrets/inside the container, out of reach ofdocker inspect. Add.envto your.gitignoreand check withgit statusthat this file is never tracked.5. Named volumes
Declare your data in named volumes with an explicit driver rather than anonymous bind-mounts. A named volume survives
docker compose down; onlydown -vremoves it. Document each volume to know what you are backing up, and associate it with a tested backup strategy: a backup without a restore test is unverified data.6. Log rotation
Add a
logging:block withdriver: json-fileand optionsmax-size: "10m"andmax-file: "3". Without this, a chatty container's logs fill the disk until failure. Apply it to every service. For high-volume stacks, consider thelocaldriver (automatic compression) or an external driver likelokiif you have a centralized logging infrastructure.7. Network isolation
Create named networks — a
frontendfor exposed services, abackendfor the database — and attach each service only to the networks it needs. Your database should only be reachable by the application, never from the shared default bridge. Addinternal: trueto thebackendnetwork to prevent any access from the host or the internet, even in the event of a port configuration error.8. Port binding
Behind a reverse proxy, publish on
127.0.0.1:8080:8080rather than8080:8080(which is equivalent to0.0.0.0). Otherwise the port remains accessible from the internet despite the proxy, bypassing your TLS and authentication rules. Note: Docker directly modifies iptables rules and can bypass UFW — binding to127.0.0.1is the only reliable guarantee on the Compose side.10. Start order
Use
depends_onwithcondition: service_healthyso a service waits not just for its dependency to start but to be genuinely available. This requires a healthcheck on the dependency (step 2) and eliminates startup races. For services that do not support a native healthcheck, you can usecondition: service_startedcombined with a wait script (wait-for-it.shordockerize).
Dev defaults vs production settings
Scroll the table
| Parameter | Dev default | Recommended in prod |
|---|---|---|
| restart | no | unless-stopped |
| healthcheck | absent | defined with interval and retries |
| memory | unconstrained | limit set (e.g. 512M) |
| secrets | plaintext possible | .env or Docker secrets |
| volumes | anonymous | named with driver |
| logs | no limit | max-size + max-file |
| network | default bridge | frontend / backend separated |
| ports | 0.0.0.0 | 127.0.0.1 behind proxy |
| image | latest | version or pinned digest |
| depends_on | start only | condition: service_healthy |
Advanced health checks: monitoring real service state
A well-configured health check goes beyond a simple curl: it measures the functional state of the service, not just network availability. For an HTTP API, test a lightweight business endpoint (/health or /ready) that returns 200 only if the database connection is established. For PostgreSQL, use pg_isready -U postgres rather than curl. For Redis, redis-cli ping is sufficient. Check the state of your containers with docker compose ps: the STATUS column shows healthy, unhealthy or starting. A container stuck in starting beyond start_period signals an initialization problem — consult docker compose logs service to find the cause. The start_period parameter is distinct from interval: it defines a grace window during which failures do not count toward retries, which is essential for services that need time to populate data or run schema migrations.
Docker secrets without Swarm: mounting files in /run/secrets/
Since Compose v2.24, the secrets: mechanism works in standalone mode (without Swarm). Declare a secret in the secrets: section at the root of the file, pointing to a local file (file: ./secrets/db_password.txt), then reference it in each service with secrets: [db_password]. Docker then mounts this file read-only in /run/secrets/db_password inside the container. The application reads the secret as an ordinary file (cat /run/secrets/db_password), and this content never appears in docker inspect, environment variables or logs. For zero-downtime rotation, version the secret: create db_password_v2 alongside db_password_v1, update the service to read v2, redeploy with docker compose up -d --no-deps service, then delete v1 once the deployment is validated. No service interruption, no window where two versions are simultaneously active on different containers.
Rollback: recovering from a failed deployment
Rolling back a Compose stack must be prepared before the deployment, not after the incident. The minimal strategy is to keep the previous image tagged and accessible (docker images lists local images; add an explicit myapp:stable tag before each update). To roll back: change the tag in your Compose file (or in your .env if the version is a variable), then rerun docker compose up -d. Data volumes are not affected by an image change, making an application rollback fast. If the previous version requires a reverse schema migration, prepare it before deploying the new version. For critical deployments, adopt a two-file scheme: docker-compose.yml (stable version) and docker-compose.override.yml (candidate version) — docker compose -f docker-compose.yml up -d instantly returns to the stable state, whatever the situation.
Always test your hardened stack locally before pushing to production: run docker compose config to validate the syntax, then docker compose up and simulate a reboot with docker compose restart. Verify that containers come back healthy and that data persists after a down followed by an up.
Troubleshooting: 5 common errors and their solutions
network not found: after a docker compose down, named networks are removed. If you restart and an external container tries to join that network, you get this error. Solution: declare shared networks as external: true or synchronize the restart of dependent stacks. port is already in use: another process (host nginx, an orphaned stack) occupies the port. Identify it with ss -tlnp | grep :PORT then stop it before restarting. OOM (Out Of Memory): the kernel kills a container without warning — look for OOMKilled: true in docker inspect container_id. Raise the memory limit (deploy.resources.limits.memory) or identify the leak with docker stats. permission denied on volume: the container process runs with a UID that has no access to the bind-mounted directory. Check the UID with docker compose exec service id, then adjust the host folder permissions (chown -R UID:GID /path) or add user: "UID:GID" to the service. Container restarting in a loop: docker compose logs --tail=50 service reveals the startup error. The most common causes are a missing environment variable, a missing config file or an unready dependency (solved by depends_on with service_healthy).
CVE-2026-17106 (CopyEscape): update Docker Engine
CVE-2026-17106, named CopyEscape, is a race condition in docker cp disclosed on 10 August 2026. An untrusted container can produce a malformed tar archive that follows a symlink outside the destination, resulting in the overwrite of an arbitrary file on the host — including the runc binary. The attack surface: any VPS running docker cp from a container whose content you do not control. The fix is available in Docker Engine ≥ 29.7.2 and Docker Desktop ≥ 4.86.0. Check your version with docker version and update before exposing a new service. If you cannot update immediately, avoid docker cp from untrusted containers and apply the principle of least privilege (--cap-drop ALL).
Backing up volumes without corruption
Copying files from a running PostgreSQL or MySQL volume with rsync or tar almost always produces a corrupted backup: the engine writes continuously during the copy, and data pages are captured at different checkpoints. The rule is to always run a SQL dump before the volume snapshot: pg_dump or mysqldump produce a consistent state you can archive or transfer. To automate this approach on a Docker VPS, tools like Restic and Offen Docker Backup orchestrate database freeze, SQL dump, encrypted snapshot and upload to remote storage. Document each named volume (step 5 of the checklist) and associate it with a tested restore strategy: a backup without a restore test is unverified data.
Conclusion
These ten settings turn a development Compose file into a production stack capable of surviving reboots, load spikes and unsupervised nights. The sections on advanced health checks, file-mounted secrets, rollback and common error troubleshooting complete this foundation and give you the reflexes to react quickly when something goes wrong. None of these points requires an additional tool: everything fits in the YAML you already have. Make a habit of running through this checklist before each deployment, ideally as a docker compose config review integrated into your deployment pipeline. Once these foundations are in place, you can add a reverse proxy like Traefik or Caddy, or an orchestration layer like Coolify, with full confidence.