Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Failure modes and recovery

The design is shaped by which failures roll back automatically and which do not. This page is the table of both, and then the three that are recent scars.

FailureCaught byRecovery
firewall / sshd / networking change locks you outdeploy-rs auto-rollbackautomatic
unbootable kernel / initrdGRUB generation menu (5 s at boot)pick the previous generation
a container fails to startnothing — deploy confirms reachability, not healthnixos-rebuild --rollback or fix forward
nftables reload wiped docker’s chainsnothing automatic; the symptom is the next container start failingsystemctl restart docker, and keep flushRuleset = false
a rotated sops secret didn’t reach a containernothing — the container holds a stale inodecopy-to-stable-path or env-file, see Secrets
tofu plan shows an unexpected diff on unchanged infrayou, reading the planthe state is wrong, not the infra — never apply; refresh/import, verify on the box
a fail2ban jail bans a docker/bridge addressnothing — a forward-chain reject on an internal IP downs every containersystemctl stop fail2ban; keep private ranges in ignoreIP
a deploy restarts every container at once and one racy unit exits non-zerodeploy-rs aborts — and its deactivation stops every containersee the abort that was worse than the failure
tofu apply fails with test(s) failed (400) on the tailnet policythe policy’s own tests, run by Tailscale at apply timethe narrowing is not live yet — see the bootstrap deadlock
a container image is gone from the registrynothing — the pull fails at the next start, on an image that has worked for monthsautomatic: ensure-image falls back to the nightly archive, then to restic — see the image archive
the DNSSEC chain breaksnothing here — kuma-check probes from the VPS, whose resolver may not validatedig +dnssec @1.1.1.1 hu-tao.dev SOA from off-box; until then the domain is dark for validating resolvers

The fail2ban blast radius

A forward-chain fail2ban jail has a blast radius the size of the whole box. If it ever bans an internal source it rejects all forwarded traffic, not one attacker — every container goes dark at once.

That is why the sites are rate-limited in caddy rather than banned in nftables: an in-process limiter can only throttle, it cannot take the forward plane down. See Access control.

The host’s ignoreIP includes 172.16.0.0/12, which holds every docker network here, so no jail can ban a bridge address at all — structurally, not because a regex happens not to match one. docker-mailserver’s own fail2ban carries the same exemption, for the same reason inside its own namespace.

Before adding any forward-chain jail, ask what happens when it bans 172.30.0.x.

The abort that was worse than the failure

The least intuitive row, and the widest. A failed activation does not leave the box on the previous generation running. It leaves it on the previous generation’s configuration, with nothing started.

flowchart TB
    u["nix flake update<br/>new nixpkgs"] --> r["every unit changes<br/>→ every container restarts at once"]
    r --> race["one racy unit exits non-zero<br/>(3-second transient)"]
    race --> abort["deploy-rs samples unit state,<br/>sees one failed unit, ABORTS"]
    abort --> deact["deactivation stops<br/><b>every</b> container"]
    deact --> out["15 healthy services down, mail included<br/>6 minutes"]
    deact --> lock["switch-to-configuration refuses:<br/>Could not acquire lock"]

    classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    class abort,deact,out,lock bad

On 2026-09-16 a three-second transient in one non-critical container therefore cost fifteen healthy services, mail included, for six minutes. The abort was more destructive than the failure it was responding to.

Recovery, in order:

systemctl restart docker
# then start the container units by hand — switch-to-configuration will refuse
# with "Could not acquire lock" while deploy-rs still holds it

The structural fix is to order fragile units behind a readiness gate rather than letting them race the restart. Full writeup: 2026-09-16 — flake update rollback.

tofu state recovery

State is gitignored — it holds every value tofu has ever read. A clone has none, and apply from no state builds a second server and moves DNS to it.

tofu/imports.tf prevents that declaratively: an import block per live resource, inert while state tracks it, active when it does not. tofu init && tofu plan then rebuilds state.

A correct recovery plan reads 32 to import, 0 to add, 1 to change, 0 to destroy — the one change being three provider-side booleans on hcloud_server.vps that the importer never sets. Anything else means stop.

Expect a second change, tailscale_acl.main, until the policy in this repo has been applied at least once: the import reads whatever the tailnet is currently serving, and the diff against tailscale-policy.hujson is the point. Read it before applying — that resource replaces the whole document, so anything the live policy carries and the file does not is deleted with nothing in the plan naming it as a loss. See The tailnet policy.

Three resources are in state that nothing declares — hcloud_network, hcloud_network_subnet and hcloud_server_network.vps, left over from the private network that was deleted. They cannot be given import blocks (Configuration for import target does not exist), and plans refresh them and then omit them from resource_changes entirely, which is recorded in tofu/imports.tf as not understood rather than explained away. Clearing them means tofu state rm.

The sharp edge (learned the hard way, 2026-09-05): the hcloud provider never reads public_net into state, so a post-import plan proposes adding it — and on this resource that detaches the primary IPs before reattaching. Applying it once took the mail IP off a running host. server.tf now carries lifecycle.ignore_changes = [public_net], so the block can never become an action. The primary IP is its own resource with delete protection precisely so a mistake here is minutes of downtime rather than a lost address.

A plan that disagrees with what you know to be true is a state problem, not an infrastructure problem. Never reach for apply to make it agree.

Runner-specific failures

SymptomCauseFix
jobs queue and never startthe runner cannot reach Forgejo, or its identity is wrongssh -J hutao@vps:2222 root@<runner>, then journalctl -u forgejo-runner
403 not permitted from a CI runner in a jobthe workflow hit a path outside the four-path allow-listwiden the allow-list deliberately, or change the workflow — see runner isolation
upload-artifact fails with no HTTP request at allGitHub’s action, not the Forgejo forkforgejo/upload-artifact@v5
you cannot ssh to the runnerits route in is defined by the VPS’s outbound rulesdeploy the VPS first; the runner has no tailnet and no other path
every job on nix-node fails in ~3 s, on unchanged workflowsthe daily podman system prune --all deleted the locally built job imagesystemctl restart forgejo-runner-ci-image — or just deploy; both now reload it. See the garbage collector eats it

A runner that is broken past recovery is replaced, not repaired — it holds a nix store and a job cache and nothing else. Rebuild in place, never destroy and recreate: CX instance types are limited availability.