# Four Times My Ansible Worked, and It Still Wasn't Done

Every Ansible run ends the same way: a `PLAY RECAP`, a row of numbers, and if you are lucky, `failed=0`. It is very easy to read that line as "done".

This week I automated four deployments with Ansible across AWS and Azure: a static website on two servers, a finance dashboard on an Azure VM, a Node.js bookstore backed by Amazon RDS, and an AI-assisted review process for changes before they reach a live server. Every one of those playbooks finished with `failed=0`. Every one of them turned out not to be finished.

What surprised me was that each gap was different. `failed=0` tells you one thing: the last run did not crash. It says nothing about four other questions that decide whether a playbook is actually ready:

1. **Can I prove what changed?**
2. **Does it pass everyone else's tooling, or only mine?**
3. **Can it run a second time?**
4. **Can someone review it before it runs?**

This post is about how each of those questions caught me out, what was going on underneath, and the change that fixed it. At the end there is a Definition of Done checklist you can copy.

![week9-diagram](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/23ae0a19-5c55-477e-a4b3-8c269c57491d.png)

---

## 1. The recap that proved nothing

**The project:** a multi-play playbook deploying a static website to two Ubuntu servers. Play 1 installs Nginx, Play 2 copies `index.html` into the web root and notifies a reload handler, and Play 3 runs on `localhost` and checks both servers return HTTP 200.

The first run worked. I cleared my terminal, kept going, and later captured a screenshot of the recap as my evidence of the "first deployment". When I looked at it properly, it said this:

```text
web1 : ok=5  changed=0  unreachable=0  failed=0
web2 : ok=5  changed=0  unreachable=0  failed=0
```

That was not the first deployment. A playbook that has just copied a new file and reloaded Nginx cannot report `changed=0`. I had kept the recap from a later run.

**Why it matters.** Idempotency, the property that makes Ansible worth using, can only be proven with a *pair* of runs: one that makes changes, followed by one that makes none. A single `changed=0` on its own is ambiguous. It might mean the playbook is idempotent. It might equally mean the playbook never had anything to do.

Look at the `ok` count too. It is 5 here and would be 6 on a deploying run, because a handler only runs when something notifies it. The recap was telling me what had happened. I just had not read it.

**The fix** was also the most useful Ansible habit I picked up this week. To test a playbook's claim that it owns a file, take the file away:

```bash
# Simulate drift: remove what the playbook is responsible for
ansible web -i inventory.ini -m file -a "path=/var/www/html/index.html state=absent" --become

ansible-playbook -i inventory.ini site.yml   # expect changes: the playbook repairs the drift
ansible-playbook -i inventory.ini site.yml   # expect changed=0: nothing left to repair
```

The next run did exactly what a first deployment should do:

```text
TASK [Copy index.html to the Nginx web root]
changed: [web2]
changed: [web1]

RUNNING HANDLER [Reload Nginx]
changed: [web1]
changed: [web2]

web1 : ok=6  changed=2  unreachable=0  failed=0
web2 : ok=6  changed=2  unreachable=0  failed=0
```

Then the rerun went back to `changed=0`.

![a3-07-first-run-recap](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/7890afad-05d1-4012-9ed1-6c220a6a967c.png)

![a3-08-idempotent-recap](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/b607edd3-e032-49a5-9af2-ec543def1a5d.png)

One more detail in that output: Play 1 reported `ok` for installing Nginx, even on the deploying run. Nginx had already been installed on those servers in an earlier exercise, so the `apt` module found nothing to do. That is idempotency working across projects, not a failure.

**The rule I kept:** evidence of a change is a pair of recaps, and deleting what the playbook owns is the quickest way to produce one on demand.

---

## 2. It ran on my machine. It failed the linter.

**The project:** Terraform provisions an Ubuntu VM on Azure with a network security group. Ansible then installs Nginx, Git and rsync, clones the Mini Finance site onto the server and uses `ansible.posix.synchronize` to rsync it into `/var/www/html/`, excluding `.git` and setting `www-data` ownership.

The playbook ran cleanly and Play 3 confirmed HTTP 200 with the right page content. The site loaded in a browser. Then I tried to commit, and my own pre-commit hook blocked it:

```text
syntax-check[unknown-module]: couldn't resolve module/action 'ansible.posix.synchronize'.
This often indicates a misspelling, missing collection, or incorrect module path.
mini-finance/ansible/site.yml:40:7
```

The module the linter said did not exist had just deployed my website.

**What was going on.** pre-commit does not run hooks in your environment. It builds a separate, isolated virtualenv for each hook. The ansible-lint hook installs `ansible-lint`, which pulls in `ansible-core`, and that is all. My project's virtualenv has the full `ansible` package, which bundles collections such as `ansible.posix`. Same repository, same file, two environments, two different answers.

This was the second time the same hook caught me. When I first set up the workstation, the pinned ansible-lint hook declared Python 3.14 as its language version, and my machine runs 3.12, so the hook environment failed to build at all.

**The fix** is to give the linter the same Ansible your playbooks run with, pinned to the version already in `requirements.txt`:

```yaml
- repo: https://github.com/ansible/ansible-lint
  rev: v26.9.0
  hooks:
    - id: ansible-lint
      # Upstream hook asks for python3.14; run it on this machine's Python 3.12 instead.
      language_version: python3
      # Same Ansible package as requirements.txt, so collections like ansible.posix resolve in the hook.
      additional_dependencies:
        - ansible==14.4.0
```

There was a faster "fix" available. ansible-lint can be told to treat a module as known (`mock_modules`), and the error disappears. But that also stops the linter checking that module's arguments, so a real mistake would pass silently. A useful control here: a deliberately misspelled module name, `ansible.posix.synchronise`, still fails the hook after the fix. The linter is checking modules, not skipping them.

**The rule I kept:** your linter's environment is part of your build. If CI or a teammate's hook sees a different set of dependencies from your terminal, "it works on my machine" is still true, and still not good enough.

---

## 3. The playbook that could only run once

**The project:** the EpicBook bookstore, a Node.js app using Express and Sequelize, deployed to EC2 with a private Amazon RDS for MySQL database. Three roles run in order: `common` for base packages, `nginx` as a reverse proxy from port 80 to the app on 8080, and `epicbook` to clone the app, connect it to the database, import the schema and run it under PM2.

Testing the role against a disposable server before touching AWS, the first run went through cleanly. The second run failed at the very first `git` task:

```text
msg: 'Local modifications exist in the destination: /home/.../theepicbook (force=no).'
```

**What was going on.** EpicBook keeps its database settings in `config/config.json`, and that file is tracked in the app's Git repository. My role cloned the repo and then templated the RDS credentials into that file. From Git's point of view I had edited a tracked file, so the working tree was dirty, and the `git` module rightly refuses to update a checkout with local changes. The playbook could deploy once and never again.

The obvious workarounds both break something else:

- **`force: true`** discards local changes on every run. The template then rewrites the file every run, so the task reports `changed` every run and restarts the app every run. It becomes idempotent in name only.
- **Marking the file as ignored in Git** hides the problem, and the next upstream change to that file becomes a conflict nobody sees coming.

**The fix** was to stop writing into a directory another tool owns. EpicBook's `config.json` already has a `production` entry that reads its connection from an environment variable (`JAWSDB_URL`). So the role now leaves the checkout untouched and writes a PM2 ecosystem file *outside* the repository:

```javascript
// ecosystem.config.js.j2: lives outside the Git checkout, mode 0600
module.exports = {
  apps: [{
    name: "{{ pm2_app_name }}",
    cwd: "{{ app_dir }}",
    script: "server.js",
    env: {
      NODE_ENV: "production",
      PORT: "{{ app_port }}",
      JAWSDB_URL: "mysql://{{ db_user }}:{{ db_password | urlencode }}@{{ db_host }}:{{ db_port }}/{{ db_name }}"
    }
  }]
};
```

The task that writes it uses `no_log: true` and `diff: false`, and the password reaches Ansible only through an environment variable on the controller, so it never appears in a file I commit, in Terraform output or in a diff.

A second, quieter problem showed up in the same test. The npm module reported `changed` on every run even when nothing had changed. The fix was to install dependencies only when the code changed or `node_modules` is missing:

```yaml
- name: Install the application dependencies (new code or missing node_modules only)
  community.general.npm:
    path: "{{ app_dir }}"
    production: true
  become: true
  become_user: "{{ app_user }}"
  when: epicbook_repo.changed or not epicbook_node_modules.stat.exists
```

The schema import needed the same care, because the SQL file has no `IF NOT EXISTS`. The role checks `information_schema` first and imports only when the tables are missing.

On AWS, the first run reported `ok=25 changed=15 failed=0`. The second reported `changed=0`, with the schema import, the PM2 start and the npm install all correctly skipped.

![a5-20-play-recap](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/db578f32-1edd-49a9-a991-5e45302c608f.png)

**The rule I kept:** idempotency is not a property of modules. It is a property of the whole design, including where you write files. Never write into a directory that another tool believes it owns.

---

## 4. The dry run that couldn't see

**The project:** a review step for changes to that live EpicBook server. A Bash script runs `ansible-playbook --check --diff`, collects every task that would change, sorts them into four risk categories (access and security, destructive or data changes, service disruption, routine configuration) and exits 0, 1 or 2. A Claude Code skill reads that evidence and explains it. A `PreToolUse` hook blocks any `ansible-playbook` command from Claude Code that lacks `--check`, so applying a change stays with a human.

In testing, the first dry run finished without errors. It was also blind in exactly the places that mattered most.

**What was going on.** In check mode, Ansible skips `command` and `shell` tasks, because it cannot predict what an arbitrary command would do. My playbook makes its most important decisions from read-only command tasks: *does the schema already exist?* (a `SELECT COUNT(*)` against RDS) and *is PM2 already running the app?* (`pm2 describe`). In a dry run those checks never ran, so the decisions that depended on them were never evaluated. The review could not tell me whether the next real run would write to the production database.

**The fix** was one line on each read-only task, telling Ansible that it is safe to run even during a dry run:

```yaml
- name: Check whether the EpicBook schema already exists
  ansible.builtin.command:
    argv: [mysql, "--host={{ db_host }}", "--port={{ db_port }}", "--user={{ db_user }}", --batch, --skip-column-names,
           "--execute=SELECT COUNT(*) FROM information_schema.tables WHERE table_schema='{{ db_name }}' AND table_name='Book'"]
  environment:
    MYSQL_PWD: "{{ db_password }}"
  register: epicbook_schema_check
  changed_when: false
  check_mode: false # read-only query; runs even in --check so the dry run can predict the import
```

I applied the same to `pm2 describe`, `nginx -t` and `cloud-init status --wait`. Only read-only tasks get this treatment. A `check_mode: false` on a task that changes something would turn your dry run into a real run.

With the fix, the baseline review of the live server was clean: `ok=18 changed=0`, **HEALTHY**, exit 0. Then I added a deliberately risky change, disabling root login over SSH, and the same review caught it before anything ran:

```text
Changed tasks (2):
    - common : Harden SSH by disabling root login
    - common : Restart SSH

[HIGH] Access and security
[MEDIUM] Service disruption

Overall Status:  FAILED
Exit Code:       2
```

The `--diff` output showed the exact line: `-#PermitRootLogin prohibit-password` would become `+PermitRootLogin no`.

![a6-15-risky-review](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/fa9d8e2c-46c7-4c8c-b741-f90b2a86dd5d.png)

I applied it myself, with a second SSH session open as a lifeline. The apply reported exactly the two predicted changes (`changed=2 failed=0`). I then verified access with a *fresh* SSH login, `ssh -o ControlPath=none`, because Ansible reuses open connections, and a reused connection cannot prove that new logins still work after an SSH change. The review after that was HEALTHY again.

![a6-19-ping-and-fresh-ssh](https://cdn.hashnode.com/uploads/covers/6a47e757d79bdff969e5d77f/e86ff547-3e59-46d5-90b4-5ca08024bcf0.png)

**What is still not solved.** A `command` task that *would* run still appears as `skipping` in check mode, not `changed`, and my script only counts `changed`. So a pending schema import would still slip past the summary, even though the log now shows the decision correctly. The Claude Code planning step spotted this gap. The same step also claimed the schema SQL "may drop tables". I checked, and none of the SQL files contains `DROP`. That was the clearest lesson of the exercise: the AI review was sharpest when it was reading evidence a script had produced, and least reliable when it was reasoning without it.

**The rule I kept:** a dry run is only as good as the decisions it can evaluate. Make read-only checks real, and treat anything the dry run skips as unknown, not safe.

---

## The pattern

| Question | What the recap said | What was actually missing | Fix |
|---|---|---|---|
| Can I prove what changed? | `failed=0` | A pair of runs: one changing, one not | Remove what the playbook owns, rerun twice |
| Does it pass shared tooling? | `failed=0` | The same dependencies in the lint environment | Pin `ansible` as a hook dependency |
| Can it run twice? | `failed=0` | A design that never dirties the Git checkout | Configuration outside the repository |
| Can it be reviewed first? | `failed=0` | Read-only checks that run in check mode | `check_mode: false` on read-only tasks |

Four different playbooks, four different failures, and the same recap line on all of them.

---

## A Definition of Done for a playbook

This is the checklist I now run before calling any playbook finished:

- [ ] **Changes proven:** I have a recap from a run that changed what I expected, and a rerun at `changed=0`.
- [ ] **Drift tested:** after removing something the playbook owns, the next run restores it.
- [ ] **Shared tooling passes:** lint passes in the same environment pre-commit or CI uses, and a deliberate mistake still fails it.
- [ ] **No writes into owned directories:** nothing templates into a Git checkout or a package-managed file.
- [ ] **Dry-run ready:** `--check --diff` completes, and every read-only decision task runs in check mode.
- [ ] **Secrets stay secret:** no credentials in the repository, in `--diff` output or in task output (`no_log`, `diff: false`).
- [ ] **Verified from outside:** a request from outside the server (`uri`, `curl`, a browser), and a fresh SSH login after any change to access.

None of these are exotic. Each one cost me time this week because I stopped at `failed=0`.

---

*I'm documenting everything I build in the DevOps Micro Internship as I go. You can follow my graded progress on my public [DMI progress page](https://dmi.pravinmishra.com/s/gbadedata.html).*

**P.S. This post is part of the DevOps Micro Internship (DMI) with Agentic AI — Cohort 3 — by [Pravin Mishra](https://www.linkedin.com/in/pravin-mishra-aws-trainer/). My graded progress is public: [https://dmi.pravinmishra.com/s/gbadedata.html](https://dmi.pravinmishra.com/s/gbadedata.html) · Start your DevOps journey: [https://dmi.pravinmishra.com/?utm_source=student&utm_medium=ps-blog&utm_campaign=cohort3](https://dmi.pravinmishra.com/?utm_source=student&utm_medium=ps-blog&utm_campaign=cohort3)**
