Separation of Repo and State

code for this post srewohs/homelab @ v0.1.0

“Paranoid is what they call people who imagine threats against their life. I have threats against my life.”

Garak, Star Trek: Deep Space Nine

New here? The About page covers what the lab is and why this blog exists.

I wanted to share my Ansible code. The roles and playbooks that configure my homelab are the kind of thing I’d have liked to find when I started, and putting them in public is a good way to keep myself honest about writing them properly.

The problem was that the code lived in the same repo as everything that describes my actual network: the inventory, my addresses and internal domain, the vault with every secret. My understanding is that this is normal for an Ansible repo, but it’s also why it can’t go public.

Copies don’t work (for me at least)

The obvious fix is a sanitized public copy. You can strip out the addresses and secrets, push the result somewhere public, and then repeat whenever the code changes and never get lazy or forget to do this step. I wasn’t really a fan of this approach because it would mean a manual scrubbing step before every push, and two copies that would inevitably get out of sync.

I know myself well enough to know that a process depending on me remembering to do it right every time is doomed from the start.

Split by kind of content instead

So instead of two copies, the repo became two repos that hold different things:

This works largely because of how Ansible finds variables. It loads group_vars/ and host_vars/ from the directory the inventory file is in, so pointing a playbook at the private inventory brings every lab value along with it. The two repos sit next to each other:

<parent>/
  homelab/            public: roles, playbooks, examples
  homelab-private/    private: inventory.yml, group_vars/, the vault

and playbooks run from the public one:

cd homelab
ansible-playbook -i ../homelab-private/inventory.yml playbooks/deploy-traefik.yml

Because there is nothing sensitive being kept in the public repo, nothing has to be stripped before sharing or kept in sync. There is only one copy of each item, and the risk of leaking secrets or lab details is minimized. The README has the setup if you’d like to do the same. If you spot something I’ve overlooked, issues and pull requests are welcome.

What had to change in the code

Splitting the files was trivial, but the code itself had my lab baked into it in a few places.

Role defaults. Some roles had real values as defaults: my domain, my addresses, user names. Those were moved into private group_vars, leaving each role with either a neutral placeholder or nothing at all.

Where there’s no safe default, the role asserts the value is set and fails with a message saying what’s missing. In my opinion a placeholder that silently misconfigures a host is worse than a loud failure before anything actually runs. From the Traefik role’s defaults:

# Required, no defaults (the role asserts they are set):
#   traefik_acme_email     Let's Encrypt account email
#   traefik_base_domains   list of domains; each gets a wildcard certificate
#   traefik_services       list of {name, url}; routed as <name>.<domain>

ansible.cfg. Mine had personal settings in it: the inventory path, the vault password file, host key checking turned off. Ansible reads exactly one config file and doesn’t merge them, so there is no way to layer a personal file on top of a shared one.

The public ansible.cfg now holds only portable settings, and the personal ones moved to environment variables in a gitignored .envrc, loaded by direnv. The repo ships a template:

export HOMELAB_PRIVATE="$PWD/../homelab-private"
export ANSIBLE_INVENTORY="$HOMELAB_PRIVATE/inventory.yml"
export ANSIBLE_VAULT_PASSWORD_FILE="$HOME/.ansible-vault-pass"

With ANSIBLE_INVENTORY set, the -i flag goes away.

Examples. The public repo has an examples/ directory with a fake inventory and every variable and vault variable the roles expect, using documentation addresses and CHANGEME values. Copy it into your own private repo and fill it in.

Fresh history

The old repo’s history mixes code and lab data in every commit, so the code went into a brand new repo started from git init. The old repo became the private one and kept its history.

Before doing that, I ran a secret scanner over the old repo’s entire history. It found nothing, and the vault was encrypted in every commit. That still didn’t make the history safe to publish because a scanner is looking for things shaped like secrets, and my addresses, host names and domain aren’t shaped like secrets. This is just information that I have determined to be sensitive based on my risk profile. The new repo’s commits also use a GitHub noreply address, so no personal email ends up in the history.

Guardrails

A split is only as good as the next commit. It’s so easy to paste something like a real address into a README or a role default without noticing, and while the quote above might be slightly dramatic, I’d rather keep things separate, so the public repo checks every commit.

A denylist that isn’t public. A pre-commit hook checks staged files against a list of identifying strings, including addresses, host names, and domains. The list itself lives in the private repo, since publishing a list of the things you’re trying to hide would be a bit counterproductive. The hook also checks the commit’s author and committer identity, and it fails closed: if it can’t find the denylist (say, on a fresh clone without the private repo next to it), it blocks the commit instead of quietly passing.

gitleaks, twice. gitleaks runs as a pre-commit hook and again in GitHub Actions on every push and pull request over the full history. A local hook can be skipped or never installed, but CI can’t. CI can’t see the private denylist either, so it only catches generic secrets like API keys. The local hook is the only one that knows what my lab looks like.

A vault leak check. Yes, I’m wearing a tinfoil hat. No, I won’t take it off. The last check before any push sits in the private repo, because it needs the vault. It decrypts the vault in memory, then searches the public repo’s working tree and every commit in its history for each secret value. On a hit, it prints the variable name and file path, but not the value. It has a self-test mode that plants a value and makes sure it gets found, so a broken check won’t pass silently.

How I knew nothing broke (I think)

Refactoring every role is a really great way to break something and not realize it, so I needed a way to show behavior didn’t change.

Before touching anything, I ran every playbook in check mode (--check --diff) against the real hosts and saved the output as a baseline. After each step of the refactor, I ran them again and compared task by task. The self-imposed rule for this was simply: any difference from the baseline needs an explanation, or it’s a bug.

As much as I’d love to put my full faith in it, check mode can’t see everything. Things like command and shell tasks don’t run in check mode, and tasks marked no_log will hide their output. For those stragglers, a small script resolved the variables the way a real run would (role defaults, inventory, vault) and wrote a hash of each value both before and after. Matching hashes mean the same values reach those tasks without the values ever being written down.

The baseline comparison and the hashing script both live in the private repo for now, so there’s nothing to link. If they turn out to be useful to anyone else, I can move them over. Just say pretty please and tell me how amazing I am.

What’s not done

Most of the lab is still set up by hand, and those parts will come over to Ansible one at a time. The benefit of this work now is that they land somewhere I can easily share with peace of mind.

Found a mistake, or something I overlooked? Open an issue on the repo.