# Containers Don't Hide Your Assumptions. They Expose Them

![A registration form showing the error "Registration failed"](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/dfe2dbca-18e5-416d-9a7e-07be3c173a41.png)

"Registration failed."

That was the whole error. Red text above a form, no status code, no hint. The Book Review app I had just deployed with Docker Compose was up, MySQL reported healthy, and the containers were all running. Registration still did not work.

The answer had nothing to do with Docker. The frontend was calling `http://<server>:3001/users/register`. The backend only serves routes under `/api`. The request was a plain 404, and the app turned it into a friendly sentence that told me nothing.

That moment set the tone for Week 11 of the DevOps Micro Internship. Eight Docker assignments, ending with a production-style deployment of a bookstore app called EpicBook on AWS. What surprised me was not how much Docker did for me. It was how often Docker made me say something out loud that I had been quietly assuming: which port, which path, which user, what "ready" means, where the data lives. Every one of those boundaries turned out to be a place where a problem was waiting.

Here is what I found, in the order I found it.

## The week in numbers

| What | Result |
| --- | --- |
| Assignments completed | 8, plus the optional CI/CD task |
| React image, single-stage vs multi-stage | 862 MB to 95 MB (88.98% smaller) |
| EpicBook backend, baseline vs optimized | 1.73 GB to 274 MB (84.16% smaller) |
| Public ports on the final stack | 1 (port 80, the reverse proxy) |
| Recovery after stopping the database | about 12 seconds, with no manual steps |
| Security audit on the frontend container | 1 WARN before, 6 of 6 PASS after |
| CI/CD run from push to verified deploy | 1 minute 4 seconds |

## 1\. The friendly error that was really a 404

Back to that form. The one call in the app that built its own URL, in the home page's `page.js`, included `/api`. Everything else (register, login, book details, reviews) went through a small helper, `services/api.js`, that built URLs without `/api`. The repository's own README said `NEXT_PUBLIC_API_URL` should be the bare backend address, `http://host:3001`, which is exactly what I had set. So the bug was in the code, not my configuration.

The fix was one line:

```javascript
// NEXT_PUBLIC_API_URL is the backend origin; every API route lives under /api
const API_URL = `${process.env.NEXT_PUBLIC_API_URL || "http://localhost:3001"}/api`;
```

Before changing anything, I checked every call in `api.js` against the backend's route files, and every one lined up once `/api` was added. Then I fixed it in my fork and rebuilt only the frontend, so MySQL and the backend kept running.

![DevTools showing a request to port 3001 returning 200 OK with Access-Control-Allow-Origin set to the frontend origin](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/2cf2a2fd-a3e4-4315-a00b-f3f5a51f11b2.png)

After the fix, DevTools told the full story in one frame: the browser calling `:3001/api/books/1`, a `200 OK`, and the backend answering with `Access-Control-Allow-Origin: http://13.40.214.81:3000`. That header matters. The frontend lived on port 3000 and the API on 3001, which makes them different origins, so every call depended on CORS being configured correctly on the backend.

**The lesson:** a friendly error message is a summary written by someone who was not looking at your request. Open the Network tab first.

## 2\. A 10.0 hiding in the lockfile

Before I put that app on a public IP, I checked what I was about to expose. The frontend was pinned to Next.js 15.2.3. That version is affected by CVE-2025-66478, a critical remote code execution flaw in React Server Components, rated 10.0, and this app uses the App Router it affects.

The fix for that release line was 15.2.6 or later. I moved to 15.2.9, the newest 15.2 patch, rather than jumping a major version. That kept the change small enough to test: the lockfile updated cleanly, `npm ci` installed it, and the app built.

The same review found three other things in the upstream Compose setup that I did not want to copy:

*   MySQL was published on port 3306 to the whole internet.
    
*   The database password and JWT secret were written straight into `docker-compose.yml`.
    
*   `backend/.env` and 3,621 files of `backend/node_modules` were committed to Git, even though the repo's own `.gitignore` excluded them.
    

None of these would have stopped the app from running. That is exactly why they are dangerous. My version reads every secret from a server-only `.env`, and it uses `${VAR:?}` so Compose refuses to start with a blank password instead of quietly using one.

## 3\. Size is a security decision, not a vanity number

![Bar chart: React app 862 MB single-stage vs 95 MB multi-stage; EpicBook backend 1.73 GB baseline vs 274 MB optimized](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/0c115e5a-792a-45e5-bde4-e68c57d74f79.png align="center")

For the React app, I built the same code twice. The single-stage image kept Node.js, npm and every package in `node_modules`, and came to 862 MB. The multi-stage image built the app in one stage and copied only the compiled static files into Nginx: 95 MB, 88.98% smaller. Both were built on the same `node:22-alpine` image, so the difference came entirely from what shipped.

For EpicBook's backend, the gap was 1.73 GB down to 274 MB. I want to be honest about why, because it is easy to overclaim here. EpicBook has no build step, so this is not the classic "build in one stage, ship in another" win. Most of the saving came from swapping the full Debian `node:22` image for Alpine and installing only production dependencies. It still matters: the optimized image carries far fewer packages that could contain a vulnerability, and it runs as the unprivileged `node` user instead of root.

Two small things caught me out along the way.

**Docker 29 shows two size columns.** `docker images` now reports DISK USAGE and CONTENT SIZE. A percentage only means something if both numbers come from the same column, so I used disk usage throughout and quoted compressed size separately.

**My own screenshot leaked an account ID.** The first time I captured `docker images`, the list included old images tagged with a full ECR address, and that address contains an AWS account ID. The fix cost nothing: each of those tags pointed at an image that also had a local name, so `docker rmi` on the ECR tags only untagged them. I retook the screenshot. Worth checking before you publish anything from a machine you have used for real work.

## 4\. What does "healthy" actually mean?

Compose lets one service wait for another with `depends_on`. On its own, that only waits for the container to *start*, not for the application inside it to be ready. So I added health checks and used `condition: service_healthy`.

The upstream health check for MySQL was:

```yaml
test: ["CMD", "mysqladmin", "ping", "-h", "localhost"]
```

It looks fine. It is subtly wrong. The first time MySQL starts with an empty volume, its entrypoint runs a temporary server with networking switched off while it creates users and loads seed data. `-h localhost` talks over the local socket, so it can report healthy during that phase. The backend then starts, tries to connect over the network, fails, and exits.

This is the version I used:

```yaml
healthcheck:
  test: ["CMD-SHELL", "MYSQL_PWD=\"$$MYSQL_PASSWORD\" mysqladmin ping -h 127.0.0.1 -u \"$$MYSQL_USER\" --silent"]
  interval: 10s
  timeout: 5s
  retries: 10
  start_period: 30s
```

Two details are doing the work. `-h 127.0.0.1` forces TCP, so the check only passes once the real server accepts network connections, the same way the backend connects. And the `$$` means the password is read from the container's own environment at check time, so it never appears in the Compose file or in `docker inspect`.

For EpicBook I went one step further and gave the backend a `/health` route that actually asks the database:

```javascript
app.get("/health", async (req, res) => {
  try {
    await db.sequelize.authenticate();
    res.status(200).json({ status: "ok", database: "up" });
  } catch (err) {
    res.status(503).json({ status: "error", database: "down" });
  }
});
```

A health endpoint that always returns `OK` tells you the process is alive. This one tells you whether the app can do its job. That difference paid off later, when I started breaking things on purpose.

## 5\. One public door, and a proxy that would have forgotten the way

![Architecture: public user to Nginx reverse proxy on port 80, frontend and backend on private ports, MySQL on an internal network with a persistent volume, plus a self-hosted CI/CD runner](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/210cc48c-0254-44c2-ac29-2d82365e5fe8.png align="center")

The EpicBook capstone asked for a production-oriented layout, and this is where it landed. Nginx is the only container with a published port. It sends `/assets/` to the frontend and everything else (pages, `/api/`, `/health`) to the backend. The database sits on a second Docker network marked `internal`, so it has no route to the internet at all. `ss -ltn` on the VM confirmed the host listened publicly only on 22 and 80.

Because the browser loads pages, assets and API calls from the same origin, this stack needed no CORS at all. Compare that with the Book Review app, where the separate API port made CORS mandatory. Same-origin routing removes a whole category of configuration.

The subtle problem here was one I caught on paper rather than in production. The obvious Nginx config is `proxy_pass http://backend:8080;`. Nginx resolves that name once, when it starts. If the backend container restarts and Docker gives it a new IP address, Nginx keeps sending traffic to the old one until someone reloads it. That would have quietly broken the recovery test I had planned. So I made Nginx ask Docker's DNS at request time instead:

```nginx
resolver 127.0.0.11 valid=10s ipv6=off;
set $backend  http://backend:8080;
set $frontend http://frontend:8080;

location /api/   { proxy_pass $backend; }
location /assets/ { proxy_pass $frontend; }
```

Using a variable in `proxy_pass` is what forces the lookup to happen per request. It is a small change, and it is the reason the next section's recovery worked without anyone touching the proxy.

## 6\. Breaking it on purpose

The capstone asked for two controlled failures. I expected them to be a formality. One of them was not.

**Stopping the backend** went as designed. The proxy returned `502 Bad Gateway` while it was down. After `docker compose start backend`, it was healthy again in about 8 seconds and `/health` returned 200, with no proxy reload thanks to the DNS change above.

**Stopping the database** found a real weakness in the application.

![Terminal: health returns 503 with database down, a page request returns 502, the database restarts and is healthy in about 8 seconds, the backend recovers about 4 seconds later, and both return 200](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/40b8d33c-b076-4a09-81ef-7b95dbdfc3f1.png)

With MySQL stopped, `/health` did exactly what it should: `503` with `"database":"down"`, and the backend stayed up. Then I requested the home page. The backend process exited, and the proxy returned `502`. This one was predicted. A dry run of EpicBook's own code against a local MySQL, done before touching the cloud, showed the same crash: the route handlers do not catch database errors, and in Node 22 an unhandled promise rejection ends the process.

Docker's `restart: unless-stopped` is what saved it. The backend kept restarting until the database came back. MySQL was healthy about 8 seconds after I started it, and the backend about 4 seconds after that. Both the health endpoint and the home page returned 200 again. Apart from starting the database, nothing was done by hand: the backend recovered on its own.

I could have hidden that crash. I wrote it into the runbook instead, along with the proper fix: error handling in the async routes that returns a 503 rather than taking the process down. A runbook that says "everything worked perfectly" is less useful than one that tells the next person where the edges are.

The backup drill, by contrast, was satisfyingly boring. I inserted a test author (record 54), took a `mysqldump` backup (about 36 KB, checked for completion), deleted the record, restored the file, and got record 54 back. It then survived a `docker compose down` and `up` because the data lives in a named volume. Never `down -v`, which deletes the volume and everything in it.

## 7\. Root by default

The last assignment was a security audit. I ran a read-only script against my frontend container, then used a Claude Code skill to explain the results. The rule was clear: the AI could run the audit and recommend a fix, but it was not allowed to edit files or touch containers. The change had to be mine.

![Initial audit: five PASS results and one WARN because the container runs as root](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/e8dafc0f-7869-464e-8b60-4a5c0c51e93c.png)

One finding: **WARN, container user.** The standard Nginx image starts its master process as root. My frontend only serves static files, so there was no reason for it.

I switched the frontend to `nginxinc/nginx-unprivileged:1.30-alpine`, which NGINX maintains for exactly this. It runs as user 101 and listens on 8080 instead of 80, so I also updated the frontend health check and the proxy route, bumped the image tag to `1.1.0`, rebuilt only the frontend, and restarted the proxy.

![Final audit: all six checks PASS, the container now runs as user 101](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/237f9c7b-db1e-4bcf-93dc-a26e9b4e1d70.png)

Six of six. `docker compose exec frontend id` now shows `uid=101(nginx)`, and the static assets still load through the proxy.

A confession, because it is the most useful part: I got the order wrong the first time. I had already made the edits before asking Claude to explain the finding, and its answer said so ("your source files already contain the fix"). That screenshot would have shown the fix arriving before the analysis, which misses the point of the exercise. So I stashed my changes with `git stash`, asked again against the untouched container, then restored them. The AI explained and recommended. I decided and edited. That split is the whole idea.

## 8\. Shipping it without opening a door

The optional task was CI/CD, and I nearly talked myself out of it. The usual GitHub Actions pattern deploys by SSHing into the server from GitHub's machines, which use a huge, changing range of IP addresses. That means opening port 22 to the internet, which contradicts the security story everything else in this capstone tells.

The alternative is a self-hosted runner installed on the VM itself. It connects *outbound* to GitHub and pulls jobs, so no inbound port opens. There is a catch worth knowing: GitHub advises against self-hosted runners on public repositories, because a pull request can carry a modified workflow that runs on your machine. My fork is public, so I locked it down. The workflow only runs on pushes to `main` and manual triggers, never on pull requests. The repository requires approval before any outside contributor's workflow can run. And the runner comes off the VM once grading is done.

![GitHub Actions run: build, push, deploy and verify jobs all green in 1 minute 4 seconds](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/0d25c028-4d16-4040-96c1-e3ac70a8a45c.png)

The pipeline has four jobs. **Build** tags both images with the short commit SHA. **Push** sends them to Docker Hub and logs out even if the push fails. **Deploy** checks out that exact commit on the server and runs `docker compose up -d --wait` for the frontend and backend only. **Verify** checks that the running containers carry the new tag and that `/health`, a page, the API and a CSS file all return 200 through the proxy.

The first run went green in 1 minute 4 seconds, and the containers on the VM reported `gbadedata/epicbook-backend:b4c000a`, healthy.

One detail I am glad I thought about: the deploy job writes the new image tags into the server's `.env`. Without that, the next time anyone ran `docker compose` by hand on that machine, Compose would quietly fall back to the old local images and undo the deployment.

## What is still not production-grade

Being clear about the edges is part of doing this properly:

*   **EpicBook still crashes on database errors.** Docker recovers it, but the real fix belongs in the application code.
    
*   **There is no approval gate before deploys.** In a real team, I would put the deploy job behind a protected GitHub environment with a required reviewer.
    
*   **Backups live on the same VM as the database.** They need to be copied somewhere else, for example S3, to survive losing the machine.
    
*   **Everything is plain HTTP.** A real deployment needs TLS in front of the proxy.
    
*   **The Book Review frontend runs** `next dev`**.** That is how the repository was designed, but a production build would be faster and smaller.
    

## A checklist you can steal

This is the list I would now run against any Compose stack before calling it done:

*   Only the reverse proxy publishes a port, and `ss -ltn` on the host agrees.
    
*   No secret appears in `docker-compose.yml`, the image, or any log line.
    
*   Every service that others depend on has a health check that proves readiness, not just that the process started.
    
*   `depends_on` uses `condition: service_healthy`.
    
*   The proxy resolves container names at request time, so restarts do not break routing.
    
*   Every container that does not need root runs as a non-root user.
    
*   Images are tagged with a version or commit SHA, never only `latest`.
    
*   You have stopped the backend and the database on purpose, and written down what happened.
    
*   You have restored a backup, not just taken one.
    

## The thread through all of it

Looking back, very little of this week was about Docker commands. It was about the moments where a container boundary forced me to state an assumption: this port, this path, this user, this is what "ready" means. Some of those assumptions were mine, and some were baked into the code I was deploying. Each one I wrote down explicitly became something I could test, and several of those tests failed in useful ways.

My graded progress through the programme is public here: [https://dmi.pravinmishra.com/s/gbadedata.html](https://dmi.pravinmishra.com/s/gbadedata.html)

If you are about to containerise something, I hope at least one of these saves you the hour it cost me.

* * *

**P.S. This post is part of the DevOps Micro Internship (DMI) with Agentic AI — Cohort 3 — by** [**Pravin Mishra**](https://www.linkedin.com/in/pravin-mishra-aws-trainer/)**. My graded progress is public:** [**https://dmi.pravinmishra.com/s/gbadedata.html**](https://dmi.pravinmishra.com/s/gbadedata.html) **· Start your DevOps journey:** [**https://dmi.pravinmishra.com/?utm\_source=student&utm\_medium=ps-blog&utm\_campaign=cohort3**](https://dmi.pravinmishra.com/?utm_source=student&utm_medium=ps-blog&utm_campaign=cohort3)
