Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

This is the whole of hu-tao.dev: a mail server, a Forgejo instance, a music library, a search engine, two Minecraft worlds, a Discord bot, and the CI that builds them — described by one flake, on two machines.

The guiding rule, stated once so the rest follows from it: there is no state on the server that this repo does not describe. One flake builds each machine; OpenTofu creates the machines and publishes their DNS. Anything a human would otherwise have to remember to run is a systemd unit instead.

The two machines

flowchart TB
    subgraph vps["hu-tao · CX43 · fsn1"]
        direction TB
        mail["mail · git · music<br/>search · status · minecraft"]
        pages_vol[("pages_data")]
    end

    subgraph runner["forgejo-runner · CX33 · nbg1"]
        direction TB
        jobs["job containers<br/>podman, one per job"]
        cache[("nix store<br/>Actions cache")]
    end

    runner -- "HTTPS 443 only<br/>fetch jobs, clone, upload artifacts" --> vps
    vps -- "ssh 22, for deploys<br/>initiated here, never there" --> runner
    vps -. "pulls published artifacts<br/>on each run, hourly as a safety net" .-> pages_vol

    classDef trusted fill:#1f6f43,stroke:#0d3a23,color:#fff
    classDef hostile fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    class vps trusted
    class runner hostile

The green box holds everything. The red box is treated as hostile — it runs code from pull requests, so the design assumes a job will one day escape its container and get root there. What that costs is bounded by the split: a rooted runner holds a nix store, a job cache and its own runner token, and can reach exactly one thing on the VPS, over HTTPS, through an allow-list that permits four URL paths. See CI runner isolation.

Both directions matter and only one is symmetric-looking. The runner reaches git.hu-tao.dev the way any client on the internet does. The VPS reaches the runner over ssh because it is the deploy host. There is no private network between them, deliberately — there was one, it had a single member, and it was deleted.

Where to start

If you want toRead
understand how a packet reaches a serviceNetwork and trust boundaries
know what runs here and how it is reachedThe service stack
know what a tailnet peer may reachThe tailnet policy
ship a changeDeploying
fix something that is broken nowFailure modes and recovery
install a machine from nothingProvisioning a machine
know why a thing is the way it isthe module comment — every modules/**.nix opens with why it exists and which mistake it guards against

That last row is not a deflection. This book is the index to those comments, not a replacement for them: the reasoning lives next to the code it constrains, where it cannot rot independently of it.

Conventions used here

  • A “gotcha” has a date. Anything described as learned the hard way names when, so a reader can tell a live constraint from a superstition.
  • Numbers are measured, not estimated, and say where they came from.
  • Diagrams are mermaid, in the markdown. They render in this book, in the Forgejo web UI, and in a pull request that changes one — so a diagram cannot quietly go stale behind a build step.

Hosts and boot

Three configurations, two machines

flake.nix builds three systems from two builders:

ConfigurationBuilderRoot diskUsed by
nixosConfigurations.vpsmkVps/dev/vda (virtio)the local QEMU test VM (nix run .#default)
nixosConfigurations.vps-hetznermkVps/dev/sdawhat tofu/nixos-anywhere installs, and what deploy .#vps activates
nixosConfigurations.runner-hetznermkRunner/dev/sdathe CI runner box

The first two differ in exactly one attribute — the disko device — and share every module, every container, every secret. vps-hetzner is the real system; vps exists so the same closure can be booted and tested locally.

runner-hetzner is a second nixosSystem in the same flake rather than a second flake, because it shares boot, hardware, nix, security and the disk layout verbatim. What it deliberately does not share is listed in runner/configuration.nix.

Note the argument that is absent. mkVps passes sops-nix.nixosModules.sops; mkRunner does not. The runner holds no age key and can decrypt nothing in secrets.yaml, so the module would only add a unit that fails at boot. This is enforced, not just intended — the flake carries a check named runner-has-no-secrets, and it is a nix build rather than an evaluation, because a check that is only evaluated never runs its builder and its failure branch is inert.

Inputs are minimal and all follows nixpkgs (nixos-26.05): disko (disk layout), sops-nix (secrets), deploy-rs (the deploy circuit breaker).

Boot and disk

The disk is GPT with three partitions (disk-config.nix): a 1 M EF02 BIOS-boot partition, a 1 G EF00 ESP mounted at /boot, and ext4 root filling the rest. Both machines use the same layout; the runner’s 80 GB is picked up without a line changing, because the root partition is size = "100%".

Two hardware facts drive this and are not preferences:

  • GRUB, configured for BIOS and UEFI (boot.nix). Hetzner Cloud boots these VMs in legacy BIOS mode — the running machine has no /sys/firmware/efi. systemd-boot is EFI-only and would install cleanly, then leave an unbootable box on first reboot. GRUB embeds its core image in the EF02 partition (BIOS chain-loads it) and writes /EFI/BOOT/BOOTX64.EFI (the fallback UEFI firmware looks for with no NVRAM entry), so the same closure boots either way. canTouchEfiVariables = false, because there is no efivarfs in BIOS mode.
  • virtio kernel modules in the initrd (hardware.nix). NixOS’s default availableKernelModules targets bare metal and contains no virtio drivers. Hetzner presents the disk over virtio, so without these the initrd cannot see /dev/sda, cannot mount root, and drops to an emergency shell — while the provider still reports the server running. The nixos-anywhere --vm-test cannot catch this; the test harness injects its own virtio modules.

The ESP is 1 G, not 512 M, because it is /boot and holds a kernel + initrd per generation. It cannot be grown without a reinstall, and filling it breaks the next deploy rather than the current boot.

The machines themselves

hu-tao          163906050  cx43  fsn1  167.233.24.58   mail, git, everything
                                       2a01:4f8:c015:b138::/64
forgejo-runner  166488672  cx33  nbg1  46.225.61.172   CI only

Both carry delete_protection and rebuild_protection, with user_data in lifecycle.ignore_changes. They are rebuilt in place, never recreated — CX instance types are limited availability, and a destroy/create cycle can find nothing to create.

The runners are a fleet rather than a box: for_each over runner_names in tofu/, with the address and identity maps keyed by server name rather than by index. An address on its own says nothing about which box it belongs to, and a bare list invites being correlated by index with some other list — which stays silently wrong when an entry is removed from the middle.

The VPS’s public IPv4 is its own resource with delete protection, deliberately separate from the server, so an IP handover carries mail reputation to a new box with no DNS change. See Migration off Ubuntu.

Network and trust boundaries

Traffic reaches the VPS through three doors, and each service sits behind exactly one of them.

flowchart TB
    net(("Internet"))
    tailnet(("Tailnet"))

    net --> edge["<b>Door 1</b> · Hetzner edge firewall<br/><code>tofu/modules/hetzner-firewall</code>"]
    edge --> nft["<b>Door 2</b> · host nftables<br/><code>modules/firewall.nix</code>"]
    tailnet --> pol["<b>Door 3</b> · tailnet policy file<br/><code>tofu/tailscale-policy.hujson</code>"]
    pol --> ts["tailscale0<br/>accepted wholesale on input"]

    nft --> inp["input hook"]
    nft --> fwd["prerouting DNAT → forward"]
    ts --> inp

    inp --> hostsvc["host namespace<br/>sshd :2222 · grafana :3000<br/>tempo :4317/4318 · pgbouncer :6432<br/>syncthing GUI :8384<br/>tarpit :222/2022/22222"]
    fwd --> dock["docker networks<br/><code>proxy</code> · <code>botnet</code>"]
    dock --> caddy["caddy"]
    caddy --> pub["public vhosts :443"]
    caddy --> priv["tailnet vhosts :8443"]

    classDef door fill:#2d4a7c,stroke:#16233c,color:#fff
    class edge,nft,pol door

The two firewalls

The edge firewall (tofu/) and the host firewall (modules/firewall.nix) are kept deliberately similar: each is what survives a misconfiguration of the other. A port opened at one but not the other is still closed. When something is unreachable, check both.

input vs forward — the single most load-bearing fact

A published container port is DNAT’d in prerouting and then forwarded — it never touches the input hook. Docker writes its own accepts into the ip filter table; in nftables every table’s chain runs, and an accept in docker’s table cannot rescue a packet that table inet nixos-fw drops. So:

  • A service’s published port belongs in the forward allow-list, not input. Getting it backwards yields a port the internet can reach that the firewall never authorised.
  • Host-namespace services (sshd on 2222, grafana, tempo, pgbouncer, syncthing, the SSH tarpit, prometheus) are on the input hook.
  • networking.nftables.flushRuleset must stay false. The default flushes the entire ruleset — including the tables docker owns — on every reload, and docker only rebuilds them when dockerd starts. The symptom is latent: running containers keep working, the next container start fails with iptables: No chain/target/match by that name, and recovery is systemctl restart docker.

This boundary is also why fail2ban jails for containerised services set chain_hook = forward, which carries a sharp edge of its own — see Failure modes.

Docker networks

NetworkSubnetPurpose
proxydocker’s poolcaddy resolves its upstreams here by container name over docker’s embedded DNS — no IP addresses in the Caddyfile
botnet172.30.0.0/24 (pinned)the discord bot + its redis, isolated from the proxy. Pinned because infra.botGateway (172.30.0.1) is a literal in the bot’s DATABASE_URL, its OTLP endpoint, and the firewall’s input rule

modules/containers/default.nix creates each network as a oneshot systemd unit that every container on it requires, so a container can never start onto a network that does not exist yet.

infra.dockerBridgeGateway (172.17.0.1) is a third address that matters: it is how a container reaches a service in the host’s network namespace, which is what caddy does for grafana:3000 and syncthing’s GUI:8384. It is observed, not enforced — deliberately not pinned with the daemon’s bip setting, even though pinning is what botSubnet/botGateway do for a network this repo creates itself. bip is daemon config, so setting it restarts docker.service, which stops every container. A wrong value here costs one failed container start and a rollback; pinning it costs a full container restart on every deploy that touches the line.

The input rule that makes that route work is iifname "br-*" ip saddr != 172.30.0.0/24 tcp dport { 3000, 8384, 8443 }, and the exclusion is the point. The rule above it grants the bot exactly two ports (4317, 6432) from exactly botSubnet; without the !=, this one immediately handed the same bridge three more, so a narrow grant was followed by a broad one and only the broad one meant anything. The bot is the right container to subtract first — it is the only one here whose input is arbitrary text from strangers that it then ships to a third-party model.

What remains matched is deliberate and worth knowing: searxng, kuma, forgejo, navidrome, dozzle, cloudflared, mailserver, webmail and pages-hook share the proxy bridge with caddy, so they are still admitted to those three ports. Both services behind them are credential-protected (grafana has a real admin login with sign-up off; syncthing’s GUI password is declared by modules/syncthing.nix through guiPasswordFile and re-applied on every deploy), so this is defence in depth, not a hole being closed.

That syncthing half used to be an assertion about the live box rather than something the deploy enforced. The config directory came across from the old host with a password already set by hand, and nothing in the repo would have noticed its absence: on fresh state syncthing comes up with a generated API key and no password at all, bound to 0.0.0.0 and reachable from every bridge in the list above. That mattered more than a login form normally does, because syncthing’s REST API is not a viewer — /rest/config/folders takes a versioning block of type external with a params.command that syncthing execs, and a folder rooted at ~/.ssh writes authorized_keys as hutao, who has passwordless sudo. The credential is now declarative, so it fails closed.

Subtracting the rest needs a positive source match, which needs a pinned subnet on proxy, which means deleting a network that already exists and detaching every container on it — a maintenance window, not an edit.

One silent consequence of adding ip saddr: the rule is now IPv4-only, where the bare version matched both families. That is free today because docker’s bridges here carry no IPv6, but turning on docker IPv6 would need an ip6 saddr != sibling or those three ports go dark over v6 with every other check still passing.

Tailscale

The tailnet interface is accepted wholesale on the input hook, which is how the tailnet-only services (grafana, tempo, dozzle, pgbouncer, syncthing GUI) are kept private — by the absence of an internet rule, not by their bind address. Several bind 0.0.0.0 and rely entirely on this.

That is also why door 3 is a policy file and not an nftables rule. nftables cannot tell tailnet peers apart: every packet off tailscale0 looks the same to it, so “which peers may reach what” is a question this layer is structurally unable to answer. tofu/tailscale-policy.hujson is the layer that can, and it narrows the VPS to ten ports for the owner account — 6432 and 4317/4318 are no longer tailnet-reachable at all. See The tailnet policy.

Tailscale SSH is off (tailscale set --ssh=false). With it on, tailscaled owned port 22 on the tailnet address, and once split DNS sent git. there on tailnet devices, git over ssh landed on Tailscale SSH instead of forgejo. Administration, deploys and the runner’s jump hop all use sshd on 2222 with an ordinary key.

Three of those services also answer by name — dozzle., grafana. and syncthing. — and that is the same mechanism wearing a hat. Caddy runs a second listener on infra.tailnetHttpsPort (8443) carrying those three vhosts, plus tailnet copies of git. and status. with their admin paths open (see split DNS); like every other private port it is published on 0.0.0.0 and kept private by being in neither allow-list.

What makes the URL portless is a nat prerouting chain:

type nat hook prerouting priority -110; policy accept;
iifname tailscale0 tcp dport 443 redirect to :8443
iifname tailscale0 tcp dport 80  redirect to :8880
iifname tailscale0 udp dport 443 redirect to :8443

The udp line is HTTP/3. The tailnet listener speaks it too (8443/udp is published), and its sites advertise h3=":443" rather than caddy’s default h3=":8443": udp 8443 sent directly over the tailnet did not get through when measured, while udp 443 through this redirect did. A browser that learned the public listener’s h3=":443" takes the same path once split DNS points git. or status. at the tailnet address. Without the redirect that traffic hung — Forgejo’s webpack chunk loads failed on it — and with docker’s rule it would reach the public listener, whose copies of those names close the admin paths.

Known quirk, unexplained: udp 8443 sent directly over the tailnet does not reach caddy, while tcp 8443 direct and udp 443 through the redirect both do. Nothing uses it — the tailnet sites advertise :443 — so it is recorded here rather than chased. docker port caddy and nft list ruleset | grep 8443 on the box are where to start if it ever matters.

Two details in those lines do all the work, and they are independent:

  • iifname tailscale0 is what keeps this off public traffic. A request arriving on the public interface never matches, so it falls through to docker’s own prerouting and reaches caddy’s ordinary :443 listener exactly as it always did. This chain cannot affect a public vhost.
  • priority -110 is ten ahead of docker’s dstnat at -100, and that ordering is the entire control. Both chains run on the same hook; the lower number runs first.
flowchart TB
    pkt["tcp dport 443"] --> iif{"which interface?"}

    iif -- "public" --> dn["docker dstnat -100<br/>ours never matched"]
    dn --> pub443["caddy :443<br/>public vhosts"] --> ok1(["200"])

    iif -- "tailscale0" --> prio{"our chain's<br/>priority?"}

    prio -- "-110, ahead ✓" --> rw["redirect 443 → 8443"]
    rw --> priv["caddy :8443<br/>tailnet vhosts"] --> ok2(["200"])

    prio -- "after -100 ✗" --> dn2["docker dstnat first"]
    dn2 --> pub443b["caddy :443<br/>public vhosts"] --> nf(["404<br/>no such site"])

    classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    classDef good fill:#1f6f43,stroke:#0d3a23,color:#fff
    class nf,dn2,pub443b bad
    class ok1,ok2,rw good

The right-hand branch is the misconfiguration, not a second real path — there is no case in which a tailnet request legitimately lands on the public listener. It is drawn because getting the priority wrong fails silently: the packet is still delivered, caddy still answers, and the only symptom is a 404 on a name that resolves, from a box that is up, over a link that works.

The names are plain A records to the box’s 100.x address (tofu/modules/cloudflare-dns). A CNAME to the node’s MagicDNS name would avoid that literal — the trap modules/containers/tempo.nix documents — but Tailscale does not publish <node>.<tailnet>.ts.net in public DNS (verified 2026-09-17: empty answers from 1.1.1.1, 9.9.9.9 and 8.8.8.8), so it resolves only on a device whose MagicDNS is active and fails silently on one where it is not. The literal is the lesser failure: it goes stale only when the machine is replaced, and it goes stale loudly. The certificate covers the names as ordinary SANs, because DNS-01 never asks whether a name resolves publicly.

Why there is no private network

The two machines had a Hetzner private network. It was deleted on 2026-09-19, and the reasoning is worth keeping because the shape recurs:

  • It had one member. A private network with a single host is a subnet, not a topology.
  • hcloud firewalls do not filter private traffic. Door 1 simply does not exist on that interface.
  • The VPS input chain accepts nine ports with no iifname qualifier, so every one of them was reachable from the private interface as readily as from the internet.

Together that is a standing bypass of two of the three doors, waiting for a second member to make it exploitable. The runner reaches the VPS over ordinary public HTTPS instead, where all three doors apply and Caddy can read the request. See CI runner isolation.

The tailnet policy

The tailnet was, until 2026-09-20, the one piece of infrastructure this repo did not describe. modules/firewall.nix accepts iifname tailscale0 wholesale — that is how every tailnet-only service is kept private, by the absence of an internet rule rather than by a bind address — so what a tailnet peer could reach was decided entirely in a web console, by a document nothing here could review, diff or roll back.

The consequence was a flat trust zone. Anything on the tailnet reached sshd on 2222, pgbouncer on 6432, tempo’s two OTLP receivers, grafana, dozzle and syncthing’s GUI. One compromised phone was the whole estate. nftables cannot tell tailnet peers apart; the policy file is the only layer that can, and it is now tofu/tailscale-policy.hujson.

One document, no partial ownership

The obvious wish is for tofu to own some rules and leave the rest to the console. The provider forecloses it: tailscale_acl “controls a tailnet’s entire policy file and not just the ACLs section within it”, and it “will completely overwrite existing policy file contents”. There is no section-level ownership, and no ignore_changes that could give it — the whole document is one string attribute.

So anything omitted from that file is deleted on apply, and the plan does not say so. The first real plan here carried a nodeAttrs block granting four devices Mullvad exit-node access plus tailnet Funnel, and a second tagOwner with its ssh rule — none of it in the first draft, all of it silently on its way out. It is carried over verbatim now, with a comment saying why each line cannot be tidied. Before any apply that replaces this resource, read the live policy in the admin console and diff it by eye.

HuJSON comments survive the apply and show up in the console, so the file is also the documentation the next person reads there.

Tags are the only way to narrow anything

This is the part that decides the shape of the whole file: Tailscale has no deny rules. Rules are purely additive. You cannot narrow a destination by adding a rule — you can only stop a wider rule from covering it.

The wider rule is the stock one, device-to-device over autogroup:self. And per Tailscale’s own documentation, “autogroup:self only applies to user-owned devices. It does not apply to tagged devices.” So the moment the VPS carries a tag it drops out of that rule, and what is left is the explicit port list.

flowchart TB
    peer(("a tailnet peer"))

    peer --> pol{"policy file"}

    pol -- "autogroup:self:*<br/>every port" --> own["your own devices<br/>laptop · desktop · phone"]
    pol -- "tag:vps<br/>ten ports" --> vps["the VPS"]
    pol -- "tag:friends-ssh:22" --> friends["a friend's machine"]

    vps --> nft["host nftables<br/>iifname tailscale0 accept"]
    nft --> svc["sshd:2222 · grafana:3000<br/>dozzle:8080 · caddy:8443"]

    gone["pgbouncer :6432<br/>tempo :4317 / :4318"]
    pol -. "no rule names these" .-> gone

    classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    class gone bad

Tagging is deliberately not in the file. Applying a tag to a device is a console action, or tailscale up --advertise-tags. So moving a machine between trust classes needs no apply, no commit and no credential — which is the part that genuinely wanted to be ad-hoc — while the rules themselves change rarely and go through review.

What is live

acls: 1git2clone@github -> autogroup:self:*
      1git2clone@github -> tag:vps:22,53,80,443,2222,3000,8080,8384,8443,8880
      1git2clone@github -> tag:friends-ssh:22
ssh:  check  -> autogroup:self   as nonroot, root
      accept -> tag:friends-ssh  as nonroot, root

Every src is the owner account rather than autogroup:member, and that is a smaller change than it looks. autogroup:self as a destination means “devices belonging to whoever opened this connection”, evaluated per connection — so member A could never reach member B’s laptop through it. What naming the account removes is the one remaining case, a second member reaching their own devices. On a tailnet with one real user that is a no-op today and a closed door the day it stops being one.

The port list on tag:vps is everything a tailnet client legitimately reaches:

PortWhat
2222the host’s own sshd — deploy-critical, deploy .#vps connects here. Drop it and deploys stop
22forgejo’s ssh, for git over the tailnet
53the split-DNS resolver: answers git. and status. with this box’s tailnet address, see below
80 / 443redirected to 8880 / 8443 by the firewall’s prerouting. The ACL sees the port the client sent, before the rewrite, so these are the ones that must be listed
3000grafana, host networking
8080dozzle direct — the path that still works when caddy is the broken part
8384syncthing’s GUI
8443 / 8880caddy’s tailnet listeners, reachable directly as well as through the redirect

Deliberately absent: 6432 and 4317/4318. pgbouncer and tempo’s OTLP receivers bind 0.0.0.0 and are reached by the discord bot over the docker bridge, not from the tailnet — the input chain has a separate rule for that, scoped to botSubnet. Nothing on the tailnet has ever needed them. A tailnet peer reaching postgres is the shape of the incident this file exists to make impossible.

tag:friends-ssh is a destination and never a source

The tag predates this file. A friend’s machine carries it for exactly one reason: so it can be SSHed into without the rest of the tailnet’s members reaching it. Tagging it took it out of its owner’s autogroup:self — the same lever tag:vps uses.

One tag for the class, not one per machine. A second friend’s laptop joins this tag rather than arriving with a tagOwner, three rules and two tests of its own that a reviewer would have to read in full to discover they grant exactly what this one already grants.

Assigning it is a console action, and the order matters the same way it does for tag:vps: a device cannot be given tag:friends-ssh until an applied policy defines it, so tofu apply comes first and Machines → the device → Edit ACL tags second. A device whose tag no rule here names has no grant at all — autogroup:self does not cover it either, because it is tagged.

Nothing on those machines has any business opening a connection into this tailnet, and since there are no deny rules, “cannot initiate” is not a rule you write. It is what you get by never naming the tag as a src. The stock policy named it implicitly, because src: ["*"] covers tagged devices; deleting that wildcard is the whole fix.

Which produced the least obvious line in the file. A Tailscale SSH rule is an extra check layered on top of ordinary ACL matching, not a substitute for it — the connection must still be permitted to reach port 22 on the destination. The old dst: ["*"] granted that for free; with the wildcard gone, tag:friends-ssh:22 has to be granted explicitly in acls or ssh to those boxes simply stops working, and the ssh section looks innocent while it does. A test asserts it.

What check actually checks

The name invites a wrong guess. check does not verify that the connecting user owns the destination — that decision was already made by src/dst matching. What it adds is freshness: the connecting user re-authenticates to their Tailscale identity in a browser, and the answer is cached for about 12 hours.

So ownership comes from autogroup:self and from naming the account in src. check is what makes a stolen but still-enrolled laptop not be enough on its own. It covers your own devices only: Tailscale SSH is off on the VPS, which is administered over sshd on 2222 — see Deploying.

The tests are the part worth having

The policy above is a claim; the tests block is that claim being checked, by Tailscale, against the real tailnet:

accept  vps:2222 · vps:443 · vps:53 · vps:8080 · tag:friends-ssh:22
deny    vps:6432 · vps:4317 · tag:friends-ssh:80

A rule that stops being true fails the apply. vps:2222 is there because it is the single most load-bearing line in the file, and tag:friends-ssh:22 because it catches exactly the “tidy-up” described above.

They run at apply time, not plan time — measured, not assumed. An earlier draft of that comment said plan. A real tofu plan then succeeded against a policy carrying both an empty vps host and impossible assertions, which means the provider diffs the string locally and ships it to the API only on apply. So a broken policy reaches tofu apply before anything objects.

The provider then reports test(s) failed (400) and swallows the detail. To see it, POST the rendered file to /api/v2/tailnet/-/acl/validate with an OAuth token — that endpoint validates what you hand it and installs nothing:

[acl test error]: address "vps:6432" (protocol "tcp"): want: Drop, got: Accept
[acl test error]: address "vps:4317" (protocol "tcp"): want: Drop, got: Accept

That message is not a bug to work around. It is the file saying the narrowing is not live yet, because the VPS is not tagged.

The bootstrap deadlock

Which it was, once, and unavoidably:

  • The policy cannot be applied while the VPS is untagged, because autogroup:self:* still covers it on every port and the deny tests say otherwise.
  • The VPS cannot be tagged while the policy is unapplied, because Tailscale refuses to assign a tag that no tagOwners entry defines — and tag:vps is defined only in the policy waiting to be applied.

Neither the console nor a tailscale_device_tags resource gets around that; it is the API’s rule, not a tooling limit. var.vps_is_tagged breaks the cycle by dropping exactly those two assertions for exactly one apply:

tofu apply -var vps_is_tagged=false   # publishes tagOwners; changes no access
#  Machines -> vps -> Edit ACL tags -> tag:vps   (now offered)
tofu apply                            # the apply that actually narrows anything

It defaults to true, so the un-narrowed policy is never what you get by accident — reaching for the weaker one has to be deliberate and visible in the command. With a saved plan the variable goes on the plan, not the apply: tofu plan -var vps_is_tagged=false -out main.tfplan.

This is a three-step sequence once in the life of a tailnet, not a routine.

The VPS joins tagged, not tagged afterwards

Tagging by hand left a gap that no plan would show: a VPS rebuilt from scratch rejoins as an ordinary user-owned device, which autogroup:self:* covers on every port. That is a silent return to the flat tailnet, visible only at the next tofu plan — and deploy .#vps does not run tofu.

So the tag is now half of the credential. modules/services.nix passes --advertise-tags=tag:vps, which is not optional: Tailscale refuses to register a device with an OAuth client secret untagged.

sops.templates."tailscale-authkey".content =
  "${config.sops.placeholder.tailscale_oauth_client_secret}?ephemeral=false&preauthorized=true";

An OAuth client secret, not a tskey-auth- key, because the auth keys it replaces capped out at 90 days — the box was one forgotten rotation away from being unable to rejoin its own tailnet after a rebuild.

The two query parameters are in the clear on purpose, and one of them is load-bearing in a way that is invisible when wrong. An OAuth-minted key is ephemeral=true by default, and an ephemeral node is removed from the tailnet when it goes offline. Left at the default, this VPS would delete itself on every reboot and come back as a new node with a new tailnet address — so the dozzle/grafana/syncthing A records and the vps host in this policy would both point at nothing. Buried inside ciphertext, that is a one-character mistake nobody can review. preauthorized=true only matters if device approval is ever turned on, and is the difference between an unattended rebuild and one that waits for a human to click approve.

Changing the flags is safe on a running node: tailscaled-autoconnect only runs tailscale up when the backend state is NeedsLogin, NeedsMachineAuth or Stopped. On a node that is already Running it exits without touching anything, so this takes effect on a fresh join and changed nothing the day it landed.

Which means the live node’s tag does not come from that flag yet, and the check for it is not the obvious one. tailscale debug prefs reads "AdvertiseTags": null on this box today, because the tag was applied in the console during the bootstrap above and the node has never re-run tailscale up since. The tag is nonetheless set, server-side, which is what the policy evaluates against. Ask the control plane, not the local prefs:

$ tailscale whois 100.109.115.12
Machine:
  Name:          vps.dikdik-cloud.ts.net
  Addresses:     [100.109.115.12/32 fd7a:115c:a1e0::493b:730d/128]
  Tags:          tag:vps

The two agree only after a rejoin. Until then --advertise-tags is insurance against the rebuild, not a description of the present state.

The OAuth client wants one scope — Keys → Auth Keys → write — and tag:vps selected, since a client can only mint keys for the tags chosen at creation. The client tofu uses for the policy itself is a different one, scoped to Policy File (write).

What this still does not cover

  • nodeAttrs is carried, not understood by this repo. The four addresses granted mullvad are devices with the integration enabled; dropping one takes that machine’s Mullvad exit nodes away with nothing in the apply output saying so. They are raw addresses because that is how the console wrote them.
  • The tailnet has one real user. Every src names that account, so a second member gets nothing until a rule says otherwise — which is the right default and also means adding a person is a policy change, not an invite.
  • nftables still accepts tailscale0 wholesale. That has not changed and should not: it is the layer that cannot tell peers apart. The policy file is now the layer in front of it, and the two are the usual arrangement here — each is what survives a misconfiguration of the other.

Split DNS for the half-public names

git.<domain> and status.<domain> are served twice: publicly, and on caddy’s tailnet listener with their admin paths open (see Access control). The tailnet-only names need nothing like this — their public A record already points at the tailnet address, which only the tailnet can reach. These two must keep their public address for everyone else, so one public record cannot serve both audiences.

Tailscale split DNS closes it: tofu’s tailscale_dns_split_nameservers sends lookups for exactly those names, from tailnet devices only, to a CoreDNS on the vps bound to its tailnet address. It answers them with that address and refuses everything else. Nothing is configured per device, phones included.

  • The name list is computed in modules/containers/caddy.nix — every host that has both a public and a tailnet copy — and tofu’s var.split_dns_subdomains must agree with it.
  • It needs port 53 in the grant above, and the DNS (write) scope on the OAuth client alongside Policy File.
  • HTTP/3 has to follow the same route. A browser that saw the public Alt-Svc: h3=":443" keeps it for 30 days and then sends QUIC to udp 443 at the tailnet address. The firewall redirects tailscale0’s udp 443 to the tailnet listener, which speaks HTTP/3 as well — see Network. Before that redirect existed, those requests hung and Forgejo’s webpack chunk loads failed.
  • If the resolver is down, your tailnet devices most likely cannot resolve those two names at all. Split DNS sends them only there and Tailscale documents no fallback. Everyone else is unaffected, and both names live on this same box anyway, so the case that matters is CoreDNS alone failing, which systemd restarts. Order matters for the same reason: deploy the resolver before tofu apply points the tailnet at it.

CI runner isolation

The runner is the one machine here that runs code this repo did not write. Every other design decision in this book protects a service from the internet; this one protects the estate from its own CI.

What it replaced

The runner used to be a container on the VPS with /var/run/docker.sock bind-mounted in. That socket is root on the machine serving mail, git, every sops secret and two Minecraft servers — so any workflow, including one opened by a dependency bot, was one docker run -v /:/host away from the entire estate.

The isolation boundary did not get stronger. It moved — from a container to a VM — and what sits inside it got much smaller.

Seven layers

flowchart TB
    subgraph R["forgejo-runner · hostile"]
        direction LR
        job["job container<br/>podman, per job"]
        rd["runner daemon"]
    end

    job --> L4
    rd --> L4["④ runner nftables<br/>out: policy-drop, VPS on 443<br/>in: ssh from the VPS only"]
    L4 --> L1["① runner cloud firewall<br/>out: 53 / 80 / 443 · udp 123"]
    L1 --> L2["② VPS cloud firewall<br/>in: tcp/443 from anywhere"]
    L2 --> L3["③ VPS nftables"]
    L3 --> L5["⑤ caddy L7 allow-list<br/>4 paths, everything else 403"]
    L5 --> cad

    subgraph V["hu-tao"]
        direction LR
        cad["caddy"] --> fj["forgejo:4242"]
    end

    classDef ctl fill:#2d4a7c,stroke:#16233c,color:#fff
    class L1,L2,L3,L4,L5 ctl
#ControlWhere
1Runner cloud firewall: in = tcp/22 from the VPS /32 only; out = 53/80/443 + udp/123 (NTP)tofu/runner-firewall.tf
2VPS cloud firewall: tcp/22 outbound to the runner /32tofu/modules/hetzner-firewall
3VPS nftables: output chain is policy-drop, one rule per runner addressmodules/firewall.nix
4Runner nftables: output policy-drop, the VPS reachable on 443 and nothing else; inbound ssh accepted from the VPS alonemodules/runner/firewall.nix
5Caddy: runner addresses restricted to four Actions paths, everything else 403modules/containers/caddy.nix
6Job: no engine socket by default; its container is created per job and destroyed aftermodules/runner/default.nix
7pages-pull: a fetched artifact reduced to files and directories, modes normalised, before caddy serves itmodules/pages-pull.nix

Every connection between the two hosts is either initiated by the VPS, or is HTTPS from the runner to git.hu-tao.dev like any other client on the internet. There is no private link, deliberately — see why.

Layer 7 is a different kind of control from the six above it, which is why it was missing for a while. Every one of those reads an address, a port or a path; none of them can read what is inside a request that the allow-list legitimately permits. The published artifact is exactly that — content the runner authors, travelling through a route layer 5 has to admit, landing in a tree caddy serves. What it strips, and why.

Who may ssh in

Layers 1 and 4 both narrow inbound ssh to the VPS’s /32, and that duplication is the point — it is the same “each survives the other’s misconfiguration” arrangement as the two firewalls on the VPS. The host rule is:

ip saddr <vps4> tcp dport 22 ct state new counter name ssh_from_vps accept
tcp dport 22 ct state new counter name ssh_blocked log prefix "DROP_ssh: " drop

It used to be a bare tcp dport 22 ct state new accept, on the reasoning that only the cloud firewall should key on the VPS’s address, since a host rule would have to survive that address changing. The rest of the file had already taken that bet: the output and forward chains key their one-way drops on the same address, there is an assertion that it is non-empty, and the VM test exists precisely to override it.

What settles it is that the two failure modes are not symmetric. A stale address in the egress rules fails open — the drop matches nothing, the runner reaches the new address on every port, and the one-way design is gone with no error anywhere. Stale here fails closed: nobody can ssh in, Hetzner’s web console still works, and the fix is one rebuild. Closed is the direction to be wrong in.

IPv4 only, matching the cloud rule, which lists a v4 /32 and no v6 source at all — so v6 ssh was never reachable through the layer above. The second line is redundant with the chain’s drop policy and exists to be named: the catch-all carries an anonymous counter, so without it a refused ssh is indistinguishable from any other dropped packet.

Verified live after the deploy, from the VPS:

$ ssh -J vps root@46.225.61.172 nft list counter inet nixos-fw ssh_from_vps
counter ssh_from_vps {
    packets 2 bytes 120
}

The one-way rule, and how it is enforced

The runner’s nftables output chain is policy-drop, and the three lines that matter are ordered above every broad accept:

ip  daddr <vps4> tcp dport 443 ct state new counter name vps_allowed_out accept
ip  daddr <vps4>                counter name vps_blocked_out  log prefix "DROP_vps_out: "  drop
ip6 daddr <vps6>                counter name vps_blocked_out6 log prefix "DROP_vps_out6: " drop

Order is the whole control. nftables is first-match-wins within a chain, and the chain below these carries oifname "podman*" accept and tcp dport { 53, 80, 443 } accept. Move the drop under either one and a generic tcp dport 443 matches first, so the drop never runs — and nothing visibly breaks, because the traffic it was supposed to stop is traffic that also works. The same three lines are repeated in the forward chain, because job containers route through it rather than through output.

The v6 line drops the VPS’s whole /64 outright: nothing legitimate goes there over v6, and a runner that can reach the box on any v6 address has defeated the v4 rules.

The metadata service is the host’s alone

The forward chain carries one drop the output chain deliberately does not:

iifname "podman*" ip daddr 169.254.0.0/16 counter name metadata_blocked_fwd drop

169.254.169.254 answers this server’s own user_data, and tofu/server.tf puts the box’s Forgejo registration pair there — the per-runner entry from var.runner_identities. modules/runner/identity.nix reads it once at boot from the host’s netns, which is why the output chain allows tcp/80 and this drop leaves that path alone.

A job container is a different netns, so its packets are forwarded and land here instead. Without the rule they reach the same endpoint, and one curl in a workflow returns the uuid and secret — enough to register a second runner daemon against the instance, call FetchTask, and receive other jobs along with their tokens. That is precisely the persistence container.docker_host: "-" was chosen to deny, arriving by a route that never touches a container socket.

It matches the whole 169.254.0.0/16 rather than the single address: nothing a job does has business in link-local space, and a provider that moves its endpoint within that block does not get to reopen this quietly.

Not live on the current box. It predates the mechanism and tofu holds user_data in ignore_changes, so the endpoint returns 204 today and the identity instead sits in a 0700 root-owned file a job cannot reach. The rule is there for the next runner tofu creates, which tofu/variables.tf documents as the ordinary way to add one.

Two things test this:

  • tests/runner-firewall.nix — a three-node NixOS VM test with real packets, six subtests. The third node exists only to be not the VPS: without it the ingress half could show that ssh from the VPS is accepted, which was never the half in doubt. It asserts on the counters rather than on nc’s exit status, because no sshd is listening on the test node and a permitted connection is refused exactly like a filtered one is. It needs /dev/kvm and the VPS does not have it (shared-vCPU Hetzner instance, no nested virtualisation), so it is hand-run on a machine with KVM: nix build .#checks.x86_64-linux.runner-firewall -L.
  • runner-firewall-ordering — greps the evaluated ruleset for line order. No KVM, costs seconds, runs in CI. It catches the exact regression the VM test exists for — the VPS rules sinking below a broad podman* accept — just not with a real packet.

Measured, not assumed: 443 to the VPS returns 200; 22 and 2222 are dropped, with the drop counter moving 0 → 14.

Why layer 5 exists

Layers 1–4 are address-and-port controls. They can say “that box may reach tcp/443 here” and nothing finer — and the runner must reach 443, because that is how it fetches jobs. Without something reading the request, a rooted job gets the whole Forgejo surface: every repo it can see, the web UI, all of /api/v1.

Caddy narrows that to four paths:

/api/actions/*
/twirp/github.actions.results.api.v1.ArtifactService/*
/*/*/info/refs
/*/*/git-upload-pack

That set was measured from a real run’s access log, not guessed. No /api/v1, no web UI, and — the one worth saying out loud — no git-receive-pack: the runner can clone and cannot push. That converts the open-ended risk “a compromised runner can push to repos it built” into something a packet filter could never express, because push and fetch share a port and a TLS session.

/api/actions_pipeline/* was in this list and was deliberately removed. It is the v3 artifact API, added defensively on a guess that uploads rode it; a full run’s capture never touched it. An allow-list entry nothing uses is surface, and this is the layer whose whole job is to have less of it.

remote_ip is the TCP peer and never a header: trusted_proxies is unset and every DNS record is proxied = false, so there is nothing in front of Caddy to launder an address. A rooted runner can forge any token it holds; it cannot forge its source address.

Verified from the runner: 403 on /, /api/v1/version, /explore/repos and /user/login; allowed paths proxy through; ordinary clients still 200 everywhere; a full workflow run produced no 403 at all.

The job’s engine socket

container.docker_host is -: no engine socket is mounted into the job. Podman’s socket is root on this box, so a job holding it could plant units or read the runner’s identity pair, and that would outlive the job. Without it a compromise lasts one job. services: still work, because the runner creates service containers, not the job (a9bf861, and the comment above docker_host in modules/runner/default.nix).

Podman rather than docker, and not as a preference: docker’s nftables integration is what produced the half-working published ports and forward-chain traps documented at length in modules/firewall.nix. Jobs still get a working docker command (dockerCompat plus dockerSocket), so a workflow that shells out to docker needs no edit.

forgejo-runner 13.1.0’s daemon has no --ephemeral/--once flag, so one-job-per-runner-lifetime is not available upstream. Per-job disposability comes from the container being created and destroyed per job instead.

The job image, and the second label

Every JavaScript action — actions/checkout included — is executed by a node binary inside the job container, not by the runner. The nix label is nixos/nix, which carries nix, bash, gitMinimal, curl and coreutils and no node, so on that label uses: cannot run at all and every workflow hand-writes its checkout as a git fetch.

That is a correctness trap and not only an inconvenience. A hand-written fetch is anonymous unless the author remembers to thread the job token through it — and on a public repo an anonymous fetch works. The omission is therefore invisible on five of the six repos on this instance and fatal on the sixth: cv-template, the one private repo, failed with could not read Username for 'https://git.hu-tao.dev'. actions/checkout defaults its token input to the injected job token, so on an image with node the private-repo case needs no thought from the workflow author.

So there is a second label, and it is additive:

LabelImagenodeNotes
nixnixos/nix:2.35.2nowhat this repo’s own workflows run on
nix-nodelocalhost/forgejo-ci-nix-nodeyesnix and node; built here, not pulled
ubuntu-latestnode:22-bookwormyesno nix, and no sudo — see CI
node-22node:22-bookwormyes
alpinealpine:3.22no

nix is untouched, so vps, skavex, compress and serenity-discord-bot keep building on exactly the image they build on today; a repo opts in by changing runs-on. A broken image here cannot take CI down for four working repos.

Why it is not in a registry

force_pull is false, so the runner uses a locally present image without reaching for a registry. That lets the image be an ordinary Nix derivation (modules/runner/ci-image.nix) loaded into podman at activation — pinned by flake.lock, rebuilt only when its inputs change, with no push credential, no pull secret and no package visibility to get wrong. The label’s reference comes from config.runner.ciImageRef rather than a repeated string literal, so the label and the image it names cannot drift apart.

buildLayeredImageWithNixDb, not buildLayeredImage. The plain builder copies the store paths in but leaves /nix/var/nix/db empty, and a nix that cannot read its own database treats every path in the image as absent — nix develop then rebuilds a closure that is already sitting on the disk.

The image provides /usr/bin/env and nothing else from the FHS. dockerTools links contents into /bin and creates no /usr at all, while nixos/nix ships the usual env — so a workflow that worked on the nix label dies here on the most common shebang in the ecosystem, #!/usr/bin/env node. Every binary npm and pnpm install starts that way, which is why pnpm check on skavex failed at its first step with a message naming the interpreter rather than the script.

The garbage collector eats it

podman system prune -f --all runs daily, and --all removes every image no container references. Between jobs nothing references this one, so it is deleted like any other cold image — and every job on nix-node then dies in about three seconds on a pull of a localhost/ reference no registry can serve. Observed 2026-09-21 across cv-template, skavex and nixos-dotfiles at once, on workflows that were green hours earlier and had not changed.

Two things were wrong, and the second is the one worth carrying forward:

  • Nothing put the image back. podman-prune now carries an ExecStartPost that restarts the load unit. ExecStartPost rather than OnSuccess=, because OnSuccess only starts a unit — which is the same trap one layer up.
  • The load unit had RemainAfterExit. systemd therefore believed it active forever after its first success and would not run it again; and because the unit’s own text does not change when the image does, a deploy could not restart it either. So the deploy that installed the prune hook fixed the next deletion and left the current one in place, and every job kept failing until the unit was restarted by hand. Dropping the flag is what makes both healing paths work, since activation starts wanted-but-inactive units. It only ever bought skipping a podman load whose layers were already on disk.

Identity, and why it is not a Nix string

The runner is a declared runner: a uuid+secret pair Forgejo issues for one record, written into server.connections in config.yaml. No imperative forgejo-runner register call (deprecated upstream), no .runner state file.

But this box is a snapshot cloned into N runners, so the uuid cannot be a Nix string the way it was for the VPS’s single permanent runner — every clone would claim the same runner record, which is undefined.

sequenceDiagram
    participant H as metadata
    participant I as identity.nix
    participant D as daemon

    H->>I: user_data
    I->>H: live instance-id
    Note over I: must match, or stop
    Note over I: prepend server: to config.yaml
    I->>D: start
    Note over D: read, declare, poll

Identity arrives via Hetzner user-data and is checked against the live instance-id before the pair is read, so a cloned snapshot is inert elsewhere. The static half of config.yaml (capacity, cache, engine settings) is a Nix-checked writeText; only the genuinely per-instance part is imperative.

Cache

The Actions cache is served by the runner itself on infra.cacheProxyPort (34567), reached from job containers over the per-job podman bridge and admitted by one input rule. The port is fixed rather than random because a firewall rule cannot name a port the daemon chooses at startup.

A new box starts empty, so the first run after provisioning pays full compile. serenity-discord-bot, same commit, back to back:

jobcoldwarm
test (postgres + redis services)13m2s4m20s3.0×
build (--all-features)6m43s1m32s4.4×
build (--features "opentelemetry ai-openrouter util-download")6m34s1m32s4.3×
build ()4m48s1m19s3.6×
clippy (--all-features)5m30s1m15s4.4×
clippy (--features "opentelemetry ai-openrouter util-download")4m49s1m14s3.9×
clippy ()4m34s1m7s4.1×
fmt — no cache, control48s40s—

fmt is the control: it uses no cache and did not move, which rules out “the new box is just slower” as an explanation for the cold column.

Why the cache works when artifacts did not

Both run the same hostname test. @actions/* checks GITHUB_SERVER_URL against GITHUB.COM, *.GHE.COM, *.LOCALHOST, else “GHES” — and on a Forgejo instance that always says GHES. What each package does with the answer is opposite:

  • @actions/artifact v2+ throws GHESNotSupportedError before opening a socket. actions/upload-artifact@v4 therefore fails with zero HTTP requests — invisible in server access logs, and not fixable at any layer below the action. Use forgejo/upload-artifact (and download-artifact), whose single patch is to make that check return false.
  • @actions/cache, including via Swatinem/rust-cache@v2, uses it to select the v1 API: if (isGhes()) return 'v1'. v1 is ACTIONS_CACHE_URL + _apis/artifactcache/, which forgejo-runner’s cache proxy implements.

Same check, opposite outcome. No cache found. is the successful-but-empty branch; an unreachable cache server goes to catch/reportError instead, so that message alone never means broken plumbing.

Host ephemerality, and why it is not on

Wiping the runner on every boot was considered and rejected:

  • It conflicts with the cache. rust-cache restores arbitrary files into ~/.cargo and target/, so a cache preserved across the wipe carries poisoning through it — and a cache not preserved costs the cold column above on every single run.
  • forgejo-runner already implements GitHub-style PR-scoped cache write isolation: writes from a pull request go to refs/pull/N, reads fall back to the shared scope. A PR cannot poison the cache the base branch reads.

So the honest statement is that the host is durable and the job is disposable, and the thing that makes that acceptable is how little the host holds.

TLS, DNS and mail

One certificate

Certificates are security.acme (acme.nix): one certificate named hu-tao.dev, with every infra.certSubdomains entry as a SAN, issued over DNS-01 through Cloudflare. Because the certificate is named after the apex (not after the first domain, as certbot does), reordering the SAN list cannot silently issue a second lineage under a new name.

  • caddy does not manage certificates. It reads the acme directory read-only; each site names its files explicitly (tls fullchain.pem key.pem), which turns off caddy’s own management. NixOS names the key key.pem, not certbot’s privkey.pem.
  • DNS-01 only touches _acme-challenge TXT records, so no name on the certificate needs an A record or a reachable port 80 — which is why the apex itself can be on it, and why the tailnet-only names are ordinary SANs.
  • The cert directory is group-owned by caddy (not acme), so the unprivileged caddy container reads it by group membership rather than by CAP_DAC_OVERRIDE, which it drops. reloadServices restarts caddy and the mailserver after a renewal.
  • Every site emits HSTS. Caddy adds nothing of the sort on its own, and until it was added no vhost here carried it. What it buys is the first request: every name is https-only and caddy already redirects http→https, but that redirect is a plaintext round trip an attacker on the path can answer instead — sslstrip against mail.’s login form, say. Tailnet sites carry it too; the plaintext redirect block deliberately does not, because a browser ignores HSTS on a plaintext response. No preload: that is a submission to a list baked into browser binaries, removal takes months, and it would bind the apex and therefore names this caddy does not serve.

The ACME account contact is infra.acmeEmail, and it is ivan@hu-tao.dev — the mailserver this repo runs. It used to be ivan@hu-tao.org, a Google-hosted mailbox nothing here describes, so expiry warnings were arriving somewhere this repo knows nothing about. Changing it registers a new ACME account: lego keys its account directory by contact address, so the next renewal registers afresh rather than updating the existing registration. Issuance is unaffected. Three things follow that option — security.acme.defaults.email, dozzle’s admin record, and var.caa_iodef in tofu, which is a separate default that has to be kept equal by hand.

Who may issue — CAA

Without a CAA record, any of the ~150 CAs in the public trust stores may issue a certificate for hu-tao.dev, and a misissuance is a valid certificate for this domain in someone else’s hands. DNSSEC does not help: a certificate is not a DNS answer, so signing the zone says nothing about who may sign for its name.

The list came from the live certificates, not from acme.nix. Three issuance paths feed this zone and only one of them is this server — which is exactly the shape of mistake that makes CAA dangerous, because a record that omits a real issuer breaks renewals silently, roughly 30 days before an expiry:

CAWho uses it
letsencrypt.orglego on the VPS over DNS-01 — and Netlify, which serves the apex (75.2.60.5 / 99.83.231.61) and also issues through Let’s Encrypt
pki.googCloudflare Universal SSL. www is a proxied CNAME, so Cloudflare terminates TLS at its own edge, currently with Google Trust Services

Plus an iodef pointing at mailto:ivan@hu-tao.dev — the only way these records ever report that someone tried.

No issuewild records are published from here, deliberately: RFC 8659 makes issue govern wildcard issuance when no issuewild is present, and the tempting hardening — issuewild ";" to forbid wildcards outright — would break that Cloudflare edge certificate, because it is a wildcard.

Read the result honestly

Cloudflare publishes CAA records of its own when it is the DNS provider, so that Universal SSL stays renewable, and they never appear in a plan. Measured immediately after the apply on 2026-09-20, eight appeared alongside the three tofu owns, and the zone now answers:

issue      comodoca.com, digicert.com, letsencrypt.org, pki.goog, ssl.com
issuewild  the same five
iodef      mailto:ivan@hu-tao.dev

So issuance went from roughly 150 CAs to five, not to one. That is a real reduction and a modest one, and it is the ceiling for as long as anything in this zone is proxied — the three records tofu publishes cannot narrow past what Cloudflare adds back. The only route to the tighter list is to stop needing Universal SSL at all: un-proxy www, which today is a 301 to an apex Netlify serves anyway. That is a website decision rather than a security one.

dig CAA hu-tao.dev is the truth; the resource in tofu/ is only the part tofu owns.

DNSSEC

Signing was switched on in the Cloudflare dashboard long before tofu knew about it, and sat pending because the DS record was never published at the registrar — a signed zone with no chain to the root is an unsigned zone with extra steps. The DS is now published and the chain validates:

dig +dnssec @1.1.1.1 hu-tao.dev SOA | grep -E '^;; flags:.* ad'
dig +dnssec @8.8.8.8 hu-tao.dev SOA | grep -E '^;; flags:.* ad'

tofu output dnssec_ds prints the half a human pastes into the registrar. That half cannot be automated — Hostinger ships no OpenTofu provider for domain management — and it does not need to be: a DS is write-once for the life of the zone.

cloudflare_zone_dnssec.main is adoption only, and its lifecycle block is the point. An apply that touches the resource puts key material back in play, and a key rotation after the DS is published takes the whole domain dark for validating resolvers — mail included. The block is what stops a stray diff from becoming that action.

The new blind spot worth naming: a broken DNSSEC chain is a whole-domain outage that the monitoring cannot see. kuma-check probes from the VPS, whose resolver may not validate, so it would keep reporting green while the rest of the internet gets SERVFAIL.

Mail deliverability depends on three things agreeing

They are set in three different places, and nothing checks that they match:

flowchart TB
    ptr["<b>PTR (rDNS)</b><br/>tofu/rdns.tf<br/><code>smtp.hu-tao.dev</code>"]
    dms["<b>DMS hostname</b><br/>mailserver.nix<br/><code>smtp.hu-tao.dev</code>"]
    mx["<b>MX target</b><br/>Cloudflare<br/><code>smtp.hu-tao.dev</code>"]

    ptr <--> dms
    dms <--> mx
    mx <--> ptr

The PTR belongs to the primary IP, not the server. That is what makes an IP handover carry mail reputation to a new box with no DNS change — see Migration off Ubuntu.

  • SPF is -all and hard-codes the IP.
  • DKIM’s public half is published from tofu/ while its private half is a sops secret mounted into DMS. The two must be halves of one key, or every recipient fails the signature.
  • DMARC is p=quarantine.

SMTP/IMAP (25/465/587/993) are published directly — an MX must be reachable at the host, so none of it can sit behind cloudflared. Only the webmail is proxied, at mail..

postmaster@ and abuse@ are aliases, declared in a read-only postfix-virtual.cf mounted over the DMS config volume (mailserver.nix). Before that, postmaster@ answered 550 5.1.1, which also swallowed the box’s own root mail — and RFC 5321 requires it to be deliverable. setup alias add now fails on purpose: aliases are a git change, because an alias that exists only in the dms_config volume is state this repo does not describe.

The Cloudflare tokens

There are two, with different scopes, and confusing them produces a failure that looks like nothing is wrong: security.acme falls back to a self-signed certificate and starts its dependent services anyway, so caddy and DMS come up serving a placeholder. The check is the issuer, not the reachability:

ssh -p 2222 hutao@<host> sudo cat /var/lib/acme/hu-tao.dev/cert.pem \
  | openssl x509 -noout -issuer
# want: issuer=C=US, O=Let's Encrypt ...
# bad:  issuer=CN=minica root ca ...   <- placeholder, DNS-01 failed

Test a token directly rather than by triggering ACME — Let’s Encrypt caps failed validations at 5 per account per hostname per hour, and lego burns one per attempt:

curl -sS -H "Authorization: Bearer $TOKEN" \
  'https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev'

The service stack

Containers are virtualisation.oci-containers (docker backend), one module per service under modules/containers/. Two conventions hold across all of them (containers/default.nix):

  • Data is a named docker volume, never a bind-mounted host directory — so restic covers a new service the moment it declares a volume. See Data and backups.
  • Config comes from the Nix store, read-only (0444). A store path changes with its content, so systemd recreates the container on a config-only change — which is what the old Ansible setup’s recreate: always was working around.

Restart policy is forced to always, and logs go to the journal — never json-file, which would put unbounded logs on the disk and break the forgejo jail that reads the journal.

What runs, and how it is reached

ServiceExposure
caddy80/443 (+443/udp) for the public sites, and 8880/8443 for the tailnet-only ones — the second pair is kept private by its absence from the firewall’s allow-lists, and nftables rewrites tailscale0’s 80/443 onto it so those URLs carry no port. Built locally with the caddy-ratelimit module (caddy.withPlugins), not the stock image
cloudflaredTunnel connected, but nothing routes through it — git/music/mail/smtp are unproxied A records straight to the VPS, so caddy serves them directly
mailserverSMTP/IMAP direct on 25, 465, 587, 993 — an MX must reach the host
webmailroundcube, proxied at mail.
forgejoSSH on 22, so clone URLs need no port; HTTP via caddy at git.; site admin (/admin, /api/v1/admin) tailnet-only
navidrome127.0.0.1:4533, reached only through caddy at music.
kumaproxy network only, reached at status.; the admin socket only on the tailnet copy of that name
searxngproxy network only, reached at search.; the only public site behind basic_auth, with caddy rate_limit in front of the bcrypt
dozzledozzle. over the tailnet, and still 8080 directly — the direct port is deliberate, since this is what you open when caddy is the broken part
grafanahost networking, :3000, tailnet only; also grafana., which caddy reaches at the docker bridge address because host networking is invisible to docker’s DNS
tempohost networking, OTLP 4317/4318 bound to 0.0.0.0; kept private by the firewall’s input chain, not by the bind address
minecraft25565; RCON on loopback only (25575). Heap 1 G floor / 6 G ceiling — see below
minecraft2second world, MC 1.21.1 on the java21 image (world 1 is 26.1.2/java25) with its own mod list; 25566, reached via the _minecraft._tcp.mc2 SRV record; RCON on loopback only (25576). Heap 1 G / 4 G
serenity-bot-0nothing published; an outbound Discord gateway client, on the botnet network. tokio-console on 127.0.0.1:6669
serenity-redisbotnet only, no published port, no volume — a cache with a Discord fallback
postgresnot a container — a host service; unix socket + loopback only, never on botnet
pgbouncernot a container — a host service; 6432, reachable from botnet and the tailnet, kept private by the firewall’s input chain
syncthingnot a container — a host service; GUI on 8384, tailnet only; also syncthing., where caddy must rewrite the Host header or syncthing’s rebinding check answers 403
endlessh-gonot a container — a host service; the SSH tarpit on 222, 2022 and 22222, public on purpose; its metrics on 127.0.0.1:2112. See Observability
prometheusnot a container — a host service; loopback-only :9090, scraping the tarpit for grafana

The Forgejo Actions runner is not on this list any more: it lives on its own machine. See CI runner isolation.

Why the Minecraft heap is two numbers

itzg’s MEMORY sets -Xms and -Xmx to the same value, so the JVM commits the whole heap at startup and never gives any of it back. That is why an idle world with nobody on it sat at 4.38 G of RSS at 0–1% CPU. Both worlds therefore set INIT_MEMORY and MAX_MEMORY separately — a low floor, the same ceiling as before — rather than MEMORY.

The floor alone is not enough. G1 only uncommits at the end of a GC cycle, and an idle server triggers no GCs at all, so the heap would stay at its high-water mark forever. -XX:G1PeriodicGCInterval=300000 (JEP 346) is the other half: one concurrent cycle per five idle minutes, which is what actually returns the pages. The two are a pair — either one on its own does nothing useful.

RSS is still not heap. Metaspace, the code cache, GC structures and direct buffers live outside -Xmx, so the ceiling is not a bound on what the container reports.

Forgejo owns port 22, so the host’s sshd is on 2222, and that is the way in — Tailscale SSH is off. Keeping 22 is what lets git remotes stay portless: ssh has no service discovery — it reads no SRV record — so anything else has to be spelled out in every clone URL or every client’s ssh config.

caddy is built here, not pulled

caddy is one of three locally built images (with pages-hook and the serenity bot, below), and the only one built to add a module. It is built locally with the caddy-ratelimit module compiled in (caddy.withPlugins, wrapped in a minimal dockerTools image), because stock caddy has no rate limiting — and rate limiting is load-bearing for every site, see Access control.

It is still caddy 2.11.4 — the version tracks nixpkgs, which matches the tag the official image used — and keeps the full container hardening: non-root uid from ids.nix, read-only rootfs, --cap-drop=ALL, and tmpfs for /data, /config and /tmp.

Images are pinned where the tag moves

Most images here are repo:tag on a release tag, which is both a version and a promise that the content behind it does not change. Three are not, and they carry a digest as well (forgejo carries one too, on a release tag; the table lists only the images whose tag itself moves):

ImageWhy the tag alone says nothing
itzg/minecraft-server:java25a rolling JRE tag; a re-pull can swap the JRE under a live world
itzg/minecraft-server:java21the same, and it matters more — world 2’s mods are compiled against Java 21, and mixin/ASM on a newer JDK is the classic crash
redis:8-alpinea rolling minor tag

A digest also makes the image archive’s restore sound: a pinned pull either returns those exact bytes or fails, so falling back to an archived copy cannot silently substitute a different image.

Renovate is configured to match this. Its regex manager captures currentDigest as an optional group — without it the tag group would swallow the digest and leave the version unparseable — and the rule that keeps itzg/minecraft-server off automatic version bumps is split so that matchUpdateTypes: ["digest"] stays enabled. Before the pin, disabling that package read as caution and meant the opposite: a rolling tag has no version to bump, so there was nothing to be deliberate with, and the content moved on the next pull with no PR ever saying so. A digest PR is that missing signal.

docker.autoPrune is weekly and its flags are empty, so it is a plain docker system prune — dangling layers only, never a tagged or digest- referenced image, and never a volume. The runner’s podman prune is the one that runs --all, and what that costs is covered in CI runner isolation.

The serenity bot is built during activation

Upstream publishes no image, so serenity-bot.nix builds one from upstream’s own Dockerfile off a full-SHA fetchgit pin. The build runs in the activation script, not in its unit: that is the only phase of a switch where the old container is still serving, so the ~3½-minute compile costs no outage. stopIfChanged = false on the containers and the builder keeps them up until then, and serenity-bot-image.service only does the build at boot, when docker is not up during activation. The tag embeds the rev plus a hash of the enabled features and rustflags, so a features-only change rebuilds too.

Sharding is two numbers, shards and instances, both 1. The per-instance ranges are computed and checked exhaustively at eval time up to 16 shards. Above one instance redis stops being optional: the AI locks and rate limits are per-process without it. Until then serenity-redis is a cache with a Discord fallback: in-memory only, no volume, 256 MB LRU.

Deliberately not containers

  • postgres + pgbouncer (postgres.nix) — the first service whose data is not a docker volume, so its backup (pg_dumpall) is not optional; it is the only thing covering that data. On the host so a second service can share it without either owning the other’s volume. postgres is unix-socket and loopback only; pgbouncer (6432) is reachable from botnet and the tailnet.
  • syncthing (syncthing.nix) — a user service with state in ~/.config/syncthing and folders under ~/syncthing, paths kept byte-identical to the old box because navidrome bind-mounts ~/syncthing/Music and the node’s device ID is derived from the TLS keypair in the config dir. Tailnet-only GUI on 8384.
  • CoreDNS (caddy.nix) — answers split DNS for the half-public names, bound to the tailnet address on :53. See split DNS.
  • Renovate (renovate.nix) — a daily host timer, on the VPS rather than the CI runner so its cross-org write token never reaches an untrusted machine.
  • The Actions runner, which is now a different machine entirely.

Access control

Every service has exactly one gate, and they differ by what the service itself supports.

ServiceGate
grafanareal login from sops; anonymous-Admin off, sign-up off
dozzlebcrypt hash from sops, in a users.yml copied to a stable path
minecraft RCONpassword from sops, loopback only; one password per world (25575 / 25576)
serenity botdiscord token + AI key + db password, all sops
kumano seeding mechanism — the first visitor creates the admin account and the route then closes. Create it immediately after the first deploy.
syncthingGUI password from sops via guiPasswordFile, declared rather than inherited
pages-hookHMAC of the Forgejo system webhook; the secret is in sops and must match the hook’s Secret field
searxngno accounts at all — caddy’s basic_auth is the entire access control; the bcrypt hash is a sops secret handed to caddy via an env file

Site administration is tailnet-only

Where an app’s administration lives on paths its public side never uses, caddy answers those paths with a 404 on the public listener and serves the same name again on the tailnet listener with them open (blockedPaths in modules/containers/caddy.nix). Tailnet devices resolve those names to the tailnet address through split DNS, so the admin side works from any of them with nothing configured per device.

  • kuma — /socket.io/ is only the admin login and dashboard; the public status page loads everything from /api/status-page/*. kuma’s login is a message inside that socket, so no HTTP rate limit can count attempts at it: not exposing the socket is the only real second layer.
  • Forgejo — /admin (the site admin panel) and /api/v1/admin/* (its API: users, orgs, runner tokens, cron, system webhooks). Personal and org settings stay public. A stolen password or session can still use what the account owns, but not administer the instance from outside the tailnet. Forgejo also routes //admin and /api//v1/admin/* to the same handlers; caddy’s path matcher merges slashes, and a local replay confirmed those 404 too.

The block is its own handle, not a bare respond: caddy orders handle before respond, and the git site’s upstream sits inside handle blocks for the runner restriction, so a bare respond there would never run.

Rate limiting, every site by default

Every caddy site is rate-limited unless it sets rateLimit = null, and none does. It used to be opt-in, and only search. had opted in, which left the logins on music., mail. and git. open to guessing at whatever speed each app allowed — Forgejo allows any. Up to three zones per site, all keyed on the client IP (modules/containers/caddy.nix):

  • Site-wide — a flood brake: 300 a minute by default; git. 600; music., grafana. and syncthing. 1200; search. 120. Private ranges, this host’s public address and the CI runners are exempt, so kuma, renovate, pages-pull and CI are never throttled.
  • Login — 10 POSTs a minute to the login request, with no exemptions: Roundcube /?_task=login, Forgejo /user/login (and two-factor, forgot-password, sign-up), Navidrome /auth/login. Nothing on this host posts to a login form, so a private source here can only be a masked client, and a shared bucket failing closed is the right answer.
  • Basic auth — git. only: 30 a minute for requests carrying Authorization: Basic, which is how a password reaches Forgejo over git or the API without ever touching /user/login. Exempt like the site-wide zone, because renovate pushes with its token as a basic-auth password.

These are one layer, not the only one. Roundcube also locks an account after 3 failures a minute; Navidrome 0.64.1 throttles failed Subsonic logins itself; searxng’s bcrypt sits behind the limit. Forgejo has no throttle of its own, so two-factor auth on the account is its second layer.

The mail protocols (465, 587, 993) never pass through caddy. Their only brute-force control is docker-mailserver’s own fail2ban: six failures in a week buy a week’s ban. Its exemption list is widened to every docker network by a mounted fail2ban-jail.cf, because Roundcube’s IMAP logins arrive from a docker address — without it, six wrong webmail passwords would ban Roundcube itself and take webmail down for everyone.

searxng, the interesting one

It has no concept of a user, so authentication is the proxy’s job — and that turns out to have a second-order problem.

flowchart TB
    c(("client")) --> rl["caddy rate_limit<br/>per client IP"]
    rl -- "over limit" --> r429["429<br/>before any bcrypt"]
    rl -- "under limit" --> ba["basic_auth<br/>cost-14 bcrypt"]
    ba -- "bad" --> r401["401"]
    ba -- "good" --> sx["searxng"]
    sx -- "image_proxy<br/>many thumbnails per page" --> rl

    classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    class r429,r401 bad
  • caddy’s basic_auth on the search. site is the only thing between the instance and the internet. The credential is passed through caddy’s own {$VAR} env substitution from a sops-rendered env file — not a Nix string, because the Caddyfile is a world-readable store path and a bcrypt hash there is one anyone with a shell could crack.
  • basic_auth runs a cost-14 bcrypt on every request, and searxng’s image_proxy pulls many thumbnails per results page through caddy — so a password flood could turn bcrypt into CPU exhaustion. caddy’s rate_limit (the compiled-in module) caps hits per client IP and returns 429 before the bcrypt runs, ordered before basic_auth.

The choice of an in-process limiter over a fail2ban jail is deliberate, and it is the same lesson as the fail2ban blast radius: a misconfigured rate limiter throttles requests, it cannot take the box down.

The env file is read by docker at container start, so it re-resolves the sops generation symlink each time. A mounted template would pin a stale inode — see Secrets.

Identity model

modules/ids.nix assigns uids/gids to services that run under their own account, as base (1_000_000) + offset, from a hand-maintained append-only table.

The base clears every allocator the host uses — system users, nixbld, DynamicUser, subuid blocks — so nothing this repo assigns can ever collide with something NixOS allocated. The ceiling (2097151) is the largest uid a ustar/tar header can hold, which matters because these uids end up in restic snapshots.

The numbers are assigned, not hashed from the service name. A hash becomes immutable the moment the first file is written, and a rename then silently orphans every file the old name owned.

Read config.infra.serviceId.<name>, never a literal — so grep -rn serviceId finds every use.

Two services have ids: caddy (1) and pages-hook (2), which must own its bind-mounted /run directory without being root. A group per id is created so security.acme can chown the cert directory to caddy rather than to acme, which is what lets the container read its certificate by group membership instead of by a capability it drops.

The runner’s users are separate

modules/runner/users.nix is its own file rather than a shared import. The runner holds no age key, so it has no sops-provisioned accounts; its forgejo-runner user is a plain static system user, deliberately not DynamicUser.

The reason is ordering: the uuid+secret pair has to be written by a unit that runs before the daemon starts, and a dynamic uid does not exist until the unit it belongs to starts. There would be no stable owner to chown the composed config to ahead of time. See CI runner isolation.

Data and backups

flowchart TB
    vols[("docker volumes<br/>/var/lib/docker/volumes")]
    pg["postgres (host service)"]
    reg["container images"]
    dumps[("pg_dumpall<br/>/var/backup/postgresql")]
    imgs[("docker save<br/>/var/lib/image-archive")]
    mc[("minecraft_data<br/>minecraft2_data")]

    daily["restic · daily 00:00–01:00<br/>keep 7d / 4w / 6m"]
    weekly["restic · weekly<br/>servers STOPPED"]
    b2[("Backblaze B2")]

    pg -- "23:15, before restic" --> dumps
    reg -- "23:30, before restic" --> imgs
    dumps --> daily
    imgs --> daily
    vols --> daily
    mc -- "excluded from daily" --> weekly
    daily --> b2
    weekly --> b2

Wholesale, not per service

services.restic backs up /var/lib/docker/volumes wholesale, so a service added later is covered the moment it declares a volume. A backup that has to be told about each new service eventually stops covering one. The same rule shapes the other two paths: both are generated from what the configuration already declares, never from a hand-maintained list.

postgres runs on the host, so its data is not under docker/volumes; services.postgresqlBackup writes a pg_dumpall (every database plus globals) there at 23:15, deliberately before restic’s window, so the archived dump is never up to 23 hours stale.

The image archive

The third path is not data — it is the bytes of every container image, and it exists because a registry reference is not a guarantee that the bytes are still there.

This estate has already been bitten. forgejo 16.0.2 is unpullable: the tag still resolves on codeberg, but a platform manifest inside the index was deleted, so a rebuild from scratch cannot reach it. That was a release tag, not a rolling one, which is the part that matters — pinning a digest does not cause this and staying on a tag does not prevent it. The only thing that helps is owning a copy.

So modules/image-archive.nix runs docker save over every pullable image at 23:30, into a directory the daily restic job already ships to B2. 23:30 is the same reasoning as postgresqlBackup at 23:15: an archive written after the backup window reaches B2 a day late.

Locally built images are excluded, because a docker pull can never satisfy one and each already has a unit that produces it — caddy and pages-hook (imageFile, built by dockerTools) and serenity-bot-*. serenity-redis is not in that set: it runs a registry image and is covered like everything else.

Restoring is automatic

Each pullable container’s unit gained an ExecStartPre that resolves its image in four steps, and only reaches the last two when the ones above have failed:

already present  →  docker pull  →  the archive on disk  →  restic restore

Digest pinning is what makes that fallback sound rather than merely convenient: a pinned pull either returns those exact bytes or fails, so “pull, else restore” can never quietly substitute a different image. With a floating tag it could. See the service stack.

The restore names --path /var/lib/image-archive rather than taking a bare latest, because this repository also holds the weekly Minecraft job’s snapshots and those carry no archive at all.

Not compressed, deliberately

docker save already emits the layer blobs the way the registry stores them, so zstd measured 38 M → 38 M on redis. Worse, one compressed stream per archive would defeat restic’s content-defined deduplication — which is what makes a nightly copy of 2.6 G nearly free and lets the two Minecraft images share their common base layers.

The Minecraft worlds are a separate job

They are snapshotted with the servers stopped — a live world holds region files open and is not consistent on disk. That job stops the containers in backupPrepareCommand and restarts them from backupCleanupCommand (an ExecStopPost), so the servers return whether restic succeeded or not.

Both worlds share one job, and therefore one downtime window: a second job would mean a second stop/start cycle and a second restic run against the same repository.

Every world volume must be in this job’s paths and in the daily job’s exclude. A volume missing from the exclude list is archived hot by the daily run, which is exactly the corruption this job exists to prevent.

The password is the key

The restic password is the encryption key: lose it and every snapshot is unrecoverable. It, and the other things that deliberately live outside this repo, are catalogued in the README’s “Not in this repo” section.

What the runner holds

Nothing that is backed up, and that is the point. The runner’s state is a nix store and an Actions cache — both reconstructible, both worthless to an attacker, neither in any restic job. A runner is replaced, not restored. See CI runner isolation.

Secrets

sops + age (secrets.nix). The age private key lives at /var/lib/sops-nix/key.txt on the host and is staged there before first boot by nix run .#install — without it, activation cannot decrypt anything and the machine boots with no credentials, its own login included.

Two kinds of consumer:

  • sops.secrets.* — a decrypted file at /run/secrets/…, for things read as a file: the restic password, the DKIM key, the database password, the syncthing GUI password. The last one is owner-ed to hutao rather than left at the default 0400 root, because syncthing-init reads it as services.syncthing.user and not as root.
  • sops.templates.* — a rendered file mixing secrets with literal text, for things that want KEY=value or a whole config file: most env files here (grep -rn sops.templates modules lists all sixteen), the pgbouncer userlist, dozzle’s users.yml — and the tailnet auth key, which is the interesting one.

Every key in secrets.nix must exist in secrets.yaml, or sops-install-secrets fails during activation. This is validated at build time, so a missing key fails nix build rather than only the deploy — which is why a new secret is added to secrets.yaml before the module that reads it.

acme_email is deliberately not a secret: security.acme needs it at evaluation time, and a registration contact address is not a credential. It is infra.acmeEmail.

The tailnet key is assembled, not stored

tailscale_oauth_client_secret is an OAuth client secret, not a tskey-auth- key. Tailscale accepts one in place of an auth key and it does not expire, where the auth keys it replaces capped out at 90 days — leaving the box one forgotten rotation away from being unable to rejoin its own tailnet after a rebuild.

sops holds that string and nothing else. The two query parameters that go with it live in modules/services.nix, in the clear, on purpose:

sops.templates."tailscale-authkey".content =
  "${config.sops.placeholder.tailscale_oauth_client_secret}?ephemeral=false&preauthorized=true";

ephemeral=false is mandatory and load-bearing. An OAuth-minted key defaults to ephemeral=true, and an ephemeral node is removed from the tailnet when it goes offline — so at the default this VPS would delete itself on every reboot and rejoin as a new node with a new address, silently invalidating three DNS records and the policy file’s vps host. Inside ciphertext that is a one-character mistake nobody can review; in a template it is a line in a diff. See The tailnet policy.

A rendered template’s real path is under a generation directory, and .path only symlinks to it. Docker resolves a symlink at mount time and holds that inode forever, so a rotated secret never reaches a container that mounts the template.

flowchart TB
    subgraph broken["✗ mounted template"]
        t1["/run/secrets/rendered/x<br/>(symlink)"] --> g1["…/generation-4/x"]
        d1["docker mount"] -. "resolved once,<br/>inode pinned" .-> g1
        g2["…/generation-5/x<br/>(rotated)"]
        t1 -. "now points here" .-> g2
        d1 -.->|"never sees it"| g2
    end

    broken ~~~ ok

    subgraph ok["✓ two fixes"]
        f1["copy to a stable path first<br/><code>dozzle-users.service</code><br/><code>mailserver-dkim.service</code>"]
        f2["pass as an env file<br/>docker re-reads at container start<br/>caddy · cloudflared"]
    end

Both fixes work because they re-resolve the symlink at container start rather than pinning it at mount. The env-file form is the lighter of the two and is why searxng’s bcrypt hash reaches caddy that way — see Access control.

The runner has no secrets at all

mkRunner does not pass sops-nix.nixosModules.sops. The runner holds no age key and can decrypt nothing in secrets.yaml; the only credential on the box is its own Forgejo registration pair, delivered through Hetzner user-data.

The instance-id check does not make that pair useless elsewhere, and it is worth being precise about what it does do. The stamp lives in the staged file at /var/lib/forgejo-runner-identity/userdata and is compared against the live metadata value, so a pair staged for one box is not silently adopted by another. Forgejo knows nothing about Hetzner instance IDs — it accepts the uuid and secret from anywhere. Anyone holding that pair can register a second runner daemon against the instance and start receiving jobs.

Which is why modules/runner/firewall.nix drops link-local traffic from job containers: the same user-data that delivers the pair is readable with one curl from inside a job unless the forward chain stops it. See the metadata service is the host’s alone.

This is asserted, not merely intended: checks.runner-has-no-secrets is a nix build, because a check that is only evaluated never runs its builder.

Never decrypt to inspect

Reading secrets.yaml to “check” something is how secrets end up in a terminal, a log, or a pull request. The manifest in secrets.nix is the list of what exists; the live box is the place to verify that a secret arrived (systemctl status, the consuming service’s own health), not the ciphertext.

Observability

flowchart TB
    bot["serenity bot"] -- OTLP --> tempo["tempo :4317/4318<br/>host net"]
    tempo --> graf["grafana :3000<br/>tailnet only"]

    graf ~~~ cont
    cont["every container"] -- journald --> doz["dozzle :8080<br/>tailnet"]
    cont -- journald --> jctl["journalctl"]

    tarpit["endlessh-go<br/>:222 · :2022 · :22222"] -- "metrics :2112" --> prom["prometheus :9090<br/>loopback"]
    prom --> graf

    doz ~~~ kuma
    kuma["kuma<br/>proxy net"] --> cad["caddy"] --> status["status.hu-tao.dev"]
    check["kuma-check<br/>every 5 min"] -- "probes the PUBLIC page" --> status
    check -- ping --> hc(("healthchecks.io"))

    classDef alert fill:#8c5a1f,stroke:#4d330d,color:#fff
    class check,hc alert

tempo, grafana and prometheus use host networking and are kept private by the input chain (Network), not by their bind address. kuma publishes no port and is reached only through caddy.

The self-hosted status page paradox

kuma cannot report its own host being down. That is closed by kuma-check, an out-of-band timer that probes the public status page every five minutes and pings healthchecks.io. Healthchecks alerts on the absence of a ping, so silence becomes the alert rather than an all-clear.

The principle is worth naming once: a monitor that reports nothing when it cannot run is not a monitor.

The bot’s dashboards

Four of the dashboards seeded from modules/containers/grafana-dashboards/ into the Provisioned folder are the bot’s: overview, guild, user and DMs, linked so a guild or user row drills into its own view with the time range carried across. Every panel on them is TraceQL against tempo. The fifth is the tarpit’s. Any other dashboard lives only in grafana_data, which restic already covers.

UI edits save (allowUiUpdates), but the directory is one store path, so editing any file in it re-seeds all five and discards UI edits across the folder. Export a dashboard back into git before touching its neighbours.

Two query guards are bound to a date:

  • a second target on span.guild_id =~ "Some.*" keeps spans from before 2026-09-17, when guild_id was recorded in its Debug spelling;
  • span.attachment_urls != nil on the span tables, because attachments and links changed from string to integer on 2026-09-18, and the tempo plugin panics (HTTP 500, “No data”) on a column whose type changes mid-result.

Delete both once tempo’s retention no longer reaches those dates.

The SSH tarpit

modules/tarpit.nix runs endlessh-go on 222, 2022 and 22222, three common alternative SSH ports; the dashboard splits by port. It accepts the connection and sends a random line a second, forever, so a client waiting for the SSH version string waits for hours. It guards nothing (sshd on 2222 is key-only either way); it is there to waste bots’ time and to count them.

Its Endlessh dashboard, in the Provisioned folder, shows connections, time trapped, and a map of where they came from. The map comes from -geoip_supplier=ip-api: each new client address is looked up at ip-api.com over plain http, so bot addresses go to that third party. Its free tier allows 45 lookups a minute; a client beyond that is still trapped, just without a point on the map.

prometheus exists for this dashboard alone, on loopback :9090, with grafana’s second datasource pointed at it. Its data is a 15-day window of bot statistics and is not backed up.

A tarpit port is open in three places: modules/tarpit.nix, the input chain in modules/firewall.nix, and tofu/modules/hetzner-firewall. The last one only takes effect on tofu apply.

Where to look when something is wrong

SymptomFirst place
a site is downdozzle at :8080 directly — it is what you open when caddy is the broken part
a container will not startjournalctl -u docker-<name>
a deploy failedthe deploy output itself; then journalctl -u <unit> for the unit it named
mail is not deliveredthe issuer check in TLS, DNS and mail, then DMS’s own log
a CI job never startsjournalctl -u forgejo-runner on the runner, reached with ssh -J hutao@vps:2222
traces are missingtempo is on host networking; check the input chain admits the bot’s bridge

Deploying

Two paths here, in order of how often you’ll walk them: redeploy (constantly) and first deploy to an existing NixOS machine (once per machine). Installing onto hardware that has no NixOS at all is Provisioning.

There are two deploy targets, and one of them is only reachable through the other:

deploy .#vps              # the VPS
deploy .#runner-forgejo-runner   # the CI runner, via ProxyJump through the VPS

Everything below assumes the ssh key is loaded, because it is passphrase- protected and nothing here can prompt for it:

ssh-agent -a /tmp/hutao-agent.sock >/dev/null 2>&1
SSH_AUTH_SOCK=/tmp/hutao-agent.sock ssh-add ~/.ssh/id_ed25519
export SSH_AUTH_SOCK=/tmp/hutao-agent.sock

A Permission denied (publickey) from any command here almost always means the agent is gone, not that a key is missing on a server.


1. Redeploy — the everyday path

nix flake check          # optional; deploy builds anyway
deploy .#vps

Without deploy-rs — same result, no automatic rollback:

nix develop            # provides nixos-rebuild on a non-NixOS workstation
nixos-rebuild switch --flake .#vps-hetzner --target-host hutao@vps --use-remote-sudo

If deploy-rs itself ever becomes unavailable, note that an unresolvable flake input stops the flake evaluating at all — so that fallback needs the input removed from flake.nix first, not just a different command.

That is the whole thing. deploy-rs builds locally, pushes the closure, activates it, then waits for a fresh connection to confirm the box is still reachable. If it cannot reconnect, the machine rolls itself back to the previous generation without being asked.

What it protects and what it does not:

FailureCaught by
firewall / sshd / networking change locks you outdeploy-rs auto-rollback
unbootable kernel or initrdGRUB generation menu, 5s timeout at boot
a container fails to startnot auto-rolled back — see below

The last row is deliberate. deploy-rs confirms reachability, not service health. A crashlooping container is visible and you still have ssh, so:

ssh -p 2222 hutao@vps 'systemctl --failed; systemctl status docker-<name>'
ssh -p 2222 hutao@vps 'sudo nixos-rebuild switch --rollback'

Rolling the whole system back because one container is unhappy is usually the wrong reflex — fix it forward.

sequenceDiagram
    participant W as workstation
    participant H as host
    Note over W: build the closure locally
    W->>H: push closure, activate
    W--xH: close the connection
    W->>H: reconnect, FRESH connection
    alt reachable within confirmTimeout
        W->>H: confirm
        Note over H: new generation kept
    else unreachable
        Note over H: rolls itself back, unattended
    end

confirmTimeout is 120 s — long enough for every container to be recreated on a config change, short enough that a hung activation is not an outage. The activation timeout is 900 s, raised from 300 s for the serenity-bot image build, which runs in the activation script while the old container is still serving; serenity-bot-image.service is only the boot-time fallback. Measured on this host on 2026-09-04, cold cache including the base image pulls: 3m28s.

The runner is deployed through the VPS

The runner is deliberately off the tailnet, and inbound ssh is narrowed to the VPS /32 at both layers — its cloud firewall and its own nftables — so the only route in is a jump:

deploy .#runner-forgejo-runner    # sshOpts carry ProxyJump=hutao@vps:2222
ssh -J hutao@vps:2222 root@46.225.61.172     # by hand

Magic rollback matters more here than anywhere: a mistake in modules/runner/firewall.nix locks out the only path to the box, and the jump host cannot help with that. Deploy the VPS first and the runner last when a change touches both — the runner’s route in is defined by the VPS’s outbound rules, so a VPS deploy that has not landed yet means a runner you cannot reach.

The jump hop goes to 2222. ProxyJump=hutao@vps:2222 names the port on purpose: a bare vps lands on 22, which on the tailnet used to be Tailscale SSH and is forgejo’s now that Tailscale SSH is off. Both deploys use the VPS’s own sshd with an ordinary key, which is why 2222 is listed in the policy and called deploy-critical there.

Ports and names, so nothing surprises you

  • vps is the tailnet node name, which is why deploy.nodes.vps.hostname is a name and not an address. It is not networking.hostName (hu-tao): renaming the node in the Tailscale console breaks deploys. It survives the primary-IP handover during a migration, so the same command works before and after cutover.
  • ssh is on 2222. Port 22 belongs to forgejo, so that git clone URLs need no port. Tailscale SSH is off, so 2222 with a normal key is the only ssh in.

If a deploy fails with “lacks a signature by a trusted key”

nix.settings.trusted-users must include @wheel (it does, in modules/nix.nix). If you ever deploy to a machine that predates that setting, you cannot push to it — build on the box instead:

rsync -a --delete --exclude .git -e 'ssh -p 2222' ./ hutao@vps:nixos-image/
ssh -p 2222 hutao@vps 'cd nixos-image && sudo nixos-rebuild switch --flake .#vps-hetzner'

That is also the bootstrap for the very first deploy after an install.


2. First deploy to a machine that already runs NixOS

Same as a redeploy, with two one-time steps:

# 1. Trust the host key, or deploy-rs fails with "Host key verification failed"
#    and no way to answer the prompt.
ssh-keyscan -p 2222 -H hu-tao >> ~/.ssh/known_hosts

# 2. Confirm the box can decrypt its own secrets before relying on it.
ssh -p 2222 hutao@vps 'sudo ls /run/secrets/ | wc -l'   # expect 31

31 is 30 secrets plus the rendered/ directory; the root and user password hashes live in /run/secrets-for-users. Recompute it when a module adds a key:

nix eval --json .#nixosConfigurations.vps-hetzner.config.sops.secrets \
  --apply 's: builtins.length (builtins.filter (v: !v.neededForUsers) (builtins.attrValues s)) + 1'

Then deploy .#vps.


What makes this automatic (and what used to break it)

Every item below is now in the config. They are listed because each one, when missing, produces a machine that installs with no error and then does not work — the worst failure shape there is.

SettingWhereWithout it
boot.initrd.availableKernelModules with virtiomodules/hardware.nixNixOS’s default set is bare-metal only. The initrd cannot see /dev/sda, root never mounts, and the box sits in an emergency shell while the provider still reports it running.
GRUB, not systemd-bootmodules/boot.nixHetzner Cloud boots legacy BIOS — there is no /sys/firmware/efi. systemd-boot installs cleanly and leaves an unbootable machine.
efiInstallAsRemovable, canTouchEfiVariables = falsemodules/boot.nixThere is no efivarfs in BIOS mode; bootloader installation fails outright if it tries to write NVRAM.
time.timeZonemodules/boot.nixUnset means NixOS does not manage /etc/localtime, so docker creates a directory there and every container that bind-mounts it dies with “not a directory”.
nix.settings.trusted-users = @wheelmodules/nix.nixdeploy-rs cannot push: “lacks a signature by a trusted key”.
ssh on 2222modules/services.nixPort 22 is forgejo’s. Also needs a matching rule in the Hetzner edge firewall, which is separate from the host’s nftables.

The VM test cannot catch any of these. nixos-anywhere --flake .#vps --vm-test validates disko, GRUB and that the system boots — genuinely useful, and it is what proved GRUB-on-BIOS works. But the NixOS test harness injects its own virtio modules and its own networking, so a config that boots in the test can still be unbootable on real hardware. Treat a passing VM test as “the layout and bootloader are sane”, never as “this will boot on the server”.


Verifying a machine is actually healthy

Not “the deploy said success” — these:

ssh -p 2222 hutao@vps '
  systemctl is-system-running          # want: running
  systemctl --failed                   # want: empty
  sudo ls /run/secrets | wc -l         # want: 31
  sudo docker ps --format "{{.Names}} {{.Status}}"
  for u in caddy forgejo mailserver webmail kuma navidrome searxng pages-hook minecraft minecraft2 \
           grafana tempo dozzle cloudflared serenity-bot-0 serenity-redis; do
    echo "$u restarts=$(systemctl show -p NRestarts --value docker-$u)"
  done
  systemctl is-active postgresql pgbouncer serenity-bot-image'

Non-zero NRestarts means a crashloop that systemctl is-active will happily report as active, because systemd restarts it fast enough to look healthy.

And confirm the certificate is real rather than the self-signed placeholder that security.acme installs when issuance fails — services start either way, so nothing looks wrong until you check the issuer:

ssh -p 2222 hutao@vps 'sudo cat /var/lib/acme/hu-tao.dev/cert.pem' \
  | openssl x509 -noout -issuer -enddate
# want: issuer=C=US, O=Let's Encrypt, ...
# bad:  issuer=CN=minica root ca ...   <- placeholder, DNS-01 failed

Before triggering ACME, test the Cloudflare token directly — Let’s Encrypt caps failed validations at 5 per hour and lego spends one per attempt:

ssh -p 2222 hutao@vps 'sudo bash -c "
  T=\$(cat /run/secrets/cloudflare_api_token)
  curl -sS -H \"Authorization: Bearer \$T\" \
    https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev"'

An empty result array with success: true means the token cannot see the zone — which reads as success if you only check .success.

If deploy-rs ever disappears from GitHub

The risk is bigger than losing deploy: flake inputs are fetched from source, not from the binary cache, so an unresolvable input means the flake stops evaluating entirely and nixos-rebuild --flake fails too.

Three mitigations, in order of effort:

  1. Nothing breaks while the store path is present. A locked input already realised in /nix/store is not refetched. nix flake archive copies every input into the store on purpose, and nix-store --gc is what would remove them again.

  2. Keep a copy: nix flake archive --to file:///path/to/mirror writes all inputs somewhere you control.

  3. Cut it out — a three-part edit to flake.nix, after which option 2 above is the only deploy path:

    • delete the deploy-rs entry from inputs
    • delete deploy-rs from the outputs = { ... } argument list
    • delete the deploy and checks outputs, and deploy-rs.packages.${system}.default from the devShell

    Nothing in modules/ references it, so the machine configuration itself is unaffected.

Provisioning a machine

Installing NixOS onto hardware that has none, and the maintenance that only comes up around an install: restoring data into a fresh service, and growing the disk after a resize.

For updating a machine that already runs NixOS, see Deploying.


Bare metal — a brand new server

This is meant to be close to one command. It is, provided the machine’s quirks are already in the config — see “What makes this automatic” below.

cd tofu
cp terraform.tfvars.example terraform.tfvars && $EDITOR terraform.tfvars
tofu init
tofu plan          # READ IT. Abort on any "destroy and then create" of hcloud_server.
tofu apply

tofu creates the server with your ssh key attached at creation. It does not install NixOS — that is deliberate. Run the install separately:

nix run .#install -- root@$(tofu -chdir=tofu output -raw vps_ipv4)

tofu used to own the install through nixos-anywhere’s module. That was removed: the module declares a null_resource whose creation runs a full install, so any plan made without it already in state — a fresh clone, a lost state file, a state rm — quietly proposes reinstalling a running mail server. Infrastructure and OS installation are now separate on purpose.

Installing onto a server that already exists

One command:

nix run .#install -- root@<ip>

It verifies the age key decrypts secrets.yaml before starting, stages it into a temporary extra-files tree at 0600, and selects vps-hetzner. Extra arguments are passed through to nixos-anywhere (--debug, --build-on-remote).

Equivalent by hand, if you ever need to vary it:

mkdir -p /tmp/extra/var/lib/sops-nix
install -m 0600 ~/.sops-nix/key.txt /tmp/extra/var/lib/sops-nix/key.txt

nix run github:nix-community/nixos-anywhere -- \
  --flake .#vps-hetzner \
  --target-host root@<ip> \
  --extra-files /tmp/extra

--extra-files is not optional. Without the age key at /var/lib/sops-nix/key.txt, sops-install-secrets fails during activation and the machine boots with no credentials at all — including its own root and user passwords. Check the key decrypts before installing:

SOPS_AGE_KEY_FILE=~/.sops-nix/key.txt sops -d --extract '["email"]["postmaster"]' secrets.yaml

If the target only accepts a key you do not hold, Hetzner rescue mode is the way in — enable_rescue accepts an ssh_keys list, unlike rebuild, which re-injects whatever was attached at creation:

curl -X POST -H "Authorization: Bearer $HCLOUD_TOKEN" -H 'Content-Type: application/json' \
  -d '{"type":"linux64","ssh_keys":[<key-id>]}' \
  https://api.hetzner.cloud/v1/servers/<id>/actions/enable_rescue
curl -X POST -H "Authorization: Bearer $HCLOUD_TOKEN" \
  https://api.hetzner.cloud/v1/servers/<id>/actions/reset

Rescue is a normal Linux with the disk unmounted, which is exactly what nixos-anywhere wants.

Which configuration to install

AttrDiskUse
.#vps/dev/vdathe local QEMU VM (nix run .#default)
.#vps-hetzner/dev/sdaanything on Hetzner Cloud

They are the same closure; only the disk device differs, and boot.loader.grub.device is derived from disko so the two can never disagree. Installing .#vps on Hetzner fails at disko because /dev/vda does not exist there.


Restoring a postgres dump into a new service

services.postgresqlBackup writes a pg_dumpall to /var/backup/postgresql/all.sql.zstd nightly, and restic carries it — so the usual restore is one command:

zstd -d < /var/backup/postgresql/all.sql.zstd | sudo -u postgres psql

Seeding a service from a dump made elsewhere is different, and the ordering gets one shot. The bot runs its sqlx migrations automatically on startup, so if it reaches an empty database first, its migrations create the schema and the dump’s CREATE TABLEs then collide with it. Restore before the container’s first start.

Read the dump before running anything — two of its properties decide the commands, and guessing either one wrong fails halfway through:

head -40 ~/serenity-bot-db.sql          # pg_dump (needs a target db) or pg_dumpall (has its own CREATE DATABASE)?
grep -m5 'OWNER TO' ~/serenity-bot-db.sql   # which role does it expect to exist?

A plain-SQL dump emits ALTER TABLE … OWNER TO <role>, which hard-fails under ON_ERROR_STOP=1 if that role is absent. The cheap fix is to make the config match the dump — role and db at the top of modules/postgres.nix — rather than to rewrite the dump.

Then, for a pg_dump of a single database:

sudo systemctl stop 'docker-serenity-bot-*'
sudo -u postgres psql -c 'DROP DATABASE IF EXISTS serenity_bot;'
sudo -u postgres psql -c 'CREATE DATABASE serenity_bot OWNER serenity;'
sudo -u postgres psql -v ON_ERROR_STOP=1 -d serenity_bot -f ~/serenity-bot-db.sql
sudo systemctl start docker-serenity-bot-0

ON_ERROR_STOP=1 is not optional: without it psql reports success after skipping every statement it could not apply, which leaves a half-populated database that looks restored.

Confirm the bot treats the schema as current rather than migrating it:

journalctl -u docker-serenity-bot-0 -n 50
sudo -u postgres psql -d serenity_bot -c 'table _sqlx_migrations order by version desc limit 5;'

Growing the disk after a Hetzner resize

Resizing the volume in the Hetzner console changes the block device and nothing else. disk-config.nix declares root as size = "100%", so a fresh install fills the new disk correctly — but disko only partitions at install time, so a running machine needs this once, by hand.

The symptom is that the space is invisible rather than merely unused:

lsblk -o NAME,SIZE            # sda 152.6G, but sda3 only 75.3G
sfdisk --list-free /dev/sda   # "Unpartitioned space: 0 B"  <- lying

0 B is the tell. GPT keeps a backup header at the end of the disk, so after a resize the table still describes the old geometry — last-lba points at the old final sector, and within that table the last partition genuinely does fill the disk. sfdisk says so out loud if you read past the numbers:

GPT PMBR size mismatch (160006143 != 320004095) will be corrected by write.
The backup GPT table is not on the end of the device.

Nothing can see the free space until that header moves. Neither sgdisk nor growpart is in the system closure; pull them from the pinned nixpkgs rather than adding them permanently for a once-per-machine job:

nix build --no-link --print-out-paths 'nixpkgs#gptfdisk^out'   # sgdisk
nix build --no-link --print-out-paths 'nixpkgs#cloud-utils'    # growpart
nix copy --to ssh://hutao@vps --no-check-sigs <both paths>

Take a backup first — restic runs daily, so force one if anything since the last run matters, and save the partition table as the rollback artifact:

sudo systemctl start postgresqlBackup        # get the databases into the dump
sudo systemctl start restic-backups-b2       # then offsite
sfdisk -d /dev/sda | sudo tee /root/sda-parttable-$(date +%F).sfdisk

Then four steps, in this order:

sudo sgdisk -e /dev/sda      # 1. move the backup GPT to the true end of disk
sudo growpart /dev/sda 3     # 2. extend the LAST partition into the new space
sudo partx -u /dev/sda       # 3. make the running kernel re-read the table
sudo resize2fs /dev/sda3     # 4. grow ext4 online; safe while mounted

Step 1 is the one that is easy to skip and impossible to work around: without it, step 2 finds no free space. Step 2 only changes the partition’s end offset, so no data moves — this works because root is the last partition. Step 4 needs the resize_inode feature, which is present.

Restore path if step 1 or 2 goes wrong: sfdisk /dev/sda < /root/sda-parttable-*.sfdisk, from Hetzner rescue mode if the box will not boot.

e2fsck will look alarming afterwards, and it is lying

sudo e2fsck -fn /dev/sda3
# ... ********** WARNING: Filesystem still has errors **********

This is expected and is not evidence of damage. e2fsck cannot meaningfully check a mounted read-write filesystem: the kernel holds allocation state that has not reached the on-disk bitmaps, and -n skips journal recovery. So it reports bitmap differences, wrong free counts, and “deleted inode has zero dtime” for files unlinked while still open. Declining every fix under -n is what sets the error banner.

Prove it rather than worry about it — run it twice and compare:

sudo e2fsck -fn /dev/sda3 | grep "count wrong ("
sleep 20
sudo e2fsck -fn /dev/sda3 | grep "count wrong ("

Different numbers each run means it is tracking live writes. Identical numbers would be the thing to investigate. What actually matters is tune2fs -l /dev/sda3 | grep "Filesystem state" reporting clean.

A real check needs the filesystem offline: add fsck.mode=force fsck.repair=yes to the kernel command line for one boot, or run it from rescue mode.

Why not LVM

It would not have helped much here. Growing into new space would still need the GPT header relocated and the partition extended before pvresize, lvextend, resize2fs — three commands instead of two, for the same outcome. Where it would pay is snapshots before a risky migration, and reallocating space between volumes. Switching means a reinstall, since disko partitions only at install time, so it belongs to the next bare-metal build rather than to a resize.

CI

Two workflows, one per forge. Forgejo reads .forgejo/workflows and falls back to .github/workflows only when that directory is absent — a fallback, not a union — so the presence of .forgejo/workflows/ci.yml is what keeps Forgejo off the GitHub file. Delete it and Forgejo silently starts running a workflow written for GitHub, which is how this repo once ended up with a red run on git.hu-tao.dev.

FileRuns onJobs
.forgejo/workflows/ci.ymlthe dedicated runner boxone: check
.forgejo/workflows/pages.ymlthe dedicated runner boxone: pages — builds this book, main only
.github/workflows/ci.ymlthe GitHub mirrortwo: lint and evaluate
CheckWhat
lintpre-commit run --all-files, then gitleaks across the full history
evaluateevaluates all three nixosConfigurations, then nix flake check --no-build, then builds deploy-schema and its guard, runner-firewall-ordering, runner-has-no-secrets, pages-pull-strips-symlinks, and the minecraft-mods package

The mirror is push-only. Commit here and let it flow across; anything edited on GitHub is overwritten by the next sync, and CI can lag a push until Forgejo’s mirror job runs (Synchronize Now in the repo’s mirror settings).

Why the two files differ

Not tidiness — each difference is forced.

  • Job layout. The Forgejo file is one job; the mirror’s is two. The runner keeps a nix store between runs now that it has its own box, but the .#ci shell still has to be realised, and two jobs would pay that in parallel on capacity: 2.

  • Actions, or none at all. The Forgejo job runs on the nix label — nixos/nix, which already contains Nix — so there is no cachix/install-nix-action to run and no Nix to download per run. That image carries no node, and every JavaScript action is executed by a node binary inside the job container, so the CI file has no uses: whatsoever and does its own git fetch in place of actions/checkout.

    That is a property of this label, not of the runner. Since 2026-09-21 there is also nix-node — nix and node in one locally built image — on which uses: works normally. These two workflows have not moved to it and do not need to; see the job image for when a repo should.

    pages.yml is the exception that proves it. It needs one action — upload-artifact, which has no shell equivalent — so it puts the dev shell’s node on $GITHUB_PATH first, which is why nodejs is in the ci shell despite nothing here being a node project. Skipping that step fails the job with crun: executable file 'node' not found in $PATH before the action runs at all. skavex and hutao/compress publish the same way.

    The image is not only a saving, it is the fix: on ubuntu-latest (node:22-bookworm) install-nix-action exits 127, because the branch it takes without systemd runs sudo mkdir -p /etc/nix and that image has no sudo. The job is already root, so the sudo bought nothing to begin with.

permissions: is a GitHub-only field. Forgejo ignores it with a workflow warning, which is why the Forgejo file omits it rather than carrying a line that does nothing.

Evaluation, not a build

Both forges evaluate rather than build. It catches what actually breaks this repo — a typo’d option, a missing module argument, an infinite recursion — without asking a runner to realise a multi-gigabyte closure.

--no-build is load-bearing. deploy-rs’s deploy-activate check references the system closure, so a plain nix flake check builds the whole system, and since deploy-rs follows our nixpkgs its binary is a cache miss and is compiled from source.

Five checks and one package are built rather than evaluated, and each for a reason:

CheckWhy it must build
deploy-schemavalidates deploy.json against deploy-rs’s schema — the config is otherwise only exercised by a real deploy. deploy-schema-rejects-bad-input is its guard: feed the validator a node with no hostname and fail if that is accepted, so a validator that silently reads nothing cannot pass forever
runner-firewall-orderinggreps the evaluated nftables ruleset for rule order. A check that is only evaluated never runs its builder, so its failure branch would be inert
runner-has-no-secretssame reason — it catches the runner growing a sops-install-secrets unit, which a copy-pasted module import would do
pages-pull-strips-symlinksasserts the symlink strip and chmod sit between unzip and copy, then runs those exact lines against a hostile zip — evaluating it would never run the builder
minecraft-modsa package, not a check: it is __noChroot because it queries Modrinth, and in checks it would fail nix flake check on any sandboxed machine. The GitHub job passes --option sandbox false; the Forgejo image already has sandbox = false

runner-firewall — the real two-node VM test with real packets — is not in CI. It needs /dev/kvm, and these are shared-vCPU Hetzner instances with no nested virtualisation; qemu’s TCG software emulation was measured at roughly 5× slower just to boot two minimal nodes. It stays hand-run on a machine with KVM:

nix build .#checks.x86_64-linux.runner-firewall -L

Writing a workflow that runs here

The runner is a different machine from the VPS, with no access to its volumes and a four-path allow-list to its Forgejo instance. Four consequences:

  1. The label decides whether uses: works at all. nix has no node, so a JavaScript action cannot execute on it — the failure is crun: executable file 'node' not found in $PATH, before the action runs. Pick nix-node for a workflow that wants actions/checkout or any other action, and pick it for any private repo: a hand-written git fetch is anonymous unless the author threads the job token through it, which works on a public repo and fails on a private one. Full list of labels and what each carries: the job image.
  2. upload-artifact must be the Forgejo fork. GitHub’s bundles @actions/artifact v2, which decides a Forgejo instance is GitHub Enterprise Server and throws before opening a socket — zero HTTP requests, invisible in access logs. Use forgejo/upload-artifact@v5; a bare uses: resolves against https://code.forgejo.org, which is correct for it.
  3. Publishing writes an artifact, never a volume. See Pages.
  4. A bare uses: resolves against DEFAULT_ACTIONS_URL — which defaults to https://data.forgejo.org, a mirror of actions/* and nothing third-party. A third-party action must name its host, or it fails with remote: Not found. Pointing DEFAULT_ACTIONS_URL at github.com would fix it instance-wide, at the cost of making every bare uses: resolve to whoever holds that name on an open-registration forge.

Caching works normally — see why the cache works when artifacts did not. No cache found. is the successful-but-empty branch, not a broken cache.

Known gaps

  • Actions are pinned by moving tag rather than commit SHA.
  • pages.yml still puts the dev shell’s node on $GITHUB_PATH before its one uses: step. Moving that job to nix-node would delete the step; it has not been done, because the same file is skavex’s and the two are kept identical deliberately.
  • The nix-node image is rebuilt only when flake.lock moves, so a CVE in its node or curl waits for a flake bump.
  • The artifact pull has no size cap (--max-time bounds time, not bytes).
  • Job containers do not yet use --userns=auto.

Pages

pages.hu-tao.dev is a static site caddy serves out of one docker volume. The URL layout is the directory layout, with no rewriting anywhere:

<pages_data>/<owner>/<repo>/index.html   →   https://pages.hu-tao.dev/<owner>/<repo>/

This book is published that way, at https://pages.hu-tao.dev/hutao/vps/docs/.

The direction reversed

A publishing job used to mount the pages_data volume and write into it — that is what the old in-container runner’s one-entry valid_volumes allow-list was for. A runner on its own box cannot do that, and must not: it would be the runner reaching into the VPS, which is the one thing the split forbids.

So the job uploads an artifact and the VPS fetches it.

flowchart TB
    subgraph R["runner box"]
        job["workflow job"]
    end

    job -- "upload artifact<br/>named <b>pages</b>" --> fj

    subgraph V["hu-tao"]
        direction TB
        fj["Forgejo"]
        fj -- "POST action_run_success<br/>over the docker bridge" --> hook["pages-hook<br/>verify HMAC, touch a file"]
        hook -- "systemd .path" --> pull["pages-pull"]
        clock(["hourly timer<br/>safety net"]) --> pull
        pull -- "discover repos,<br/>fetch what changed" --> fj
        pull --> vol[("pages_data")]
        vol -- read-only --> caddy["caddy"]
    end

    caddy --> url(["pages.hu-tao.dev/&lt;owner&gt;/&lt;repo&gt;/"])

    classDef net fill:#2d4a7c,stroke:#16233c,color:#fff
    class hook net

An hourly timer starts pages-pull as well, and that is a safety net rather than the mechanism — see When it runs.

Every connection is initiated on the VPS, and in fact never leaves the host — the artifact is in Forgejo’s own storage, in a container on the same box.

The workflow runs on main only. An earlier version built on pull requests and gated the upload step with if: push && main instead, which meant the riskiest step was skipped in every rehearsal: the branch introducing the workflow went green having never once run it, and the failure landed on main at merge. Gate the workflow, not the step.

No credential. The publishing repos are public and Forgejo serves /api/v1/repos/<owner>/<repo>/actions/artifacts anonymously (verified 2026-09-19 against the live instance: 200 with a bare JSON array body). The day a private repo publishes, this needs a sops token with read:repository — and not before. Do not add one speculatively.

Not a gh-pages branch, which would be the idiomatic shape. That needs git-receive-pack, which caddy now denies to runner addresses: the allow-list carries info/refs and git-upload-pack and nothing else, so a runner can clone and cannot push. Artifacts ride the twirp ArtifactService instead, sidestepping the need entirely. See CI runner isolation.

Adding a repo

One thing: give the repo a workflow that uploads an artifact named exactly pages. That is the whole of it. Nothing is added to this repo, and no deploy is run.

pages-pull enumerates every repo on the instance through /api/v1/repos/search and publishes any that holds a live artifact by that name. Uploading it is the opt-in, the same way enabling Pages is a repo-level act rather than something the hosting provider does for you.

It was not always so. There used to be an infra.pagesRepos list in modules/options.nix, so adding a page meant a commit here and a deploy .#vps — a rebuild of the machine that serves mail, in order to publish a static site. The list existed because of a misreading of the runner split: the note said the pull side “has to be told what to look for”, when it only has to be told how to find out.

Two things fall out of the change:

  • A repo that has never published is not a special case. It has no pages artifact, so it is not discovered, and nothing is logged. Under the list this was a repo you had named in error and worth a line in the journal; now it is simply most of the instance.
  • A renamed or deleted repo stops being discovered. The old list had to be edited by hand when that happened, or the unit failed every five minutes forever — hutao/critical-forest was carried as a comment explaining exactly that. The served tree is still left in place, so caddy keeps answering that path until someone removes it.

An empty discovery is treated as an error, not as “no repos publish”. This instance always has repos, so zero means the search endpoint moved or started refusing us — and the damage would be every published site silently freezing at its current content while the unit kept exiting 0.

The artifact is untrusted content

It was built on the CI runner, the machine this estate treats as hostile. Layers 1–5 over there are address-and-port-and-path controls; an artifact is the one thing that crosses the boundary carrying content, and the caddy allow-list has to admit the ArtifactService route or publishing does not work at all.

So pages-pull reduces the unpacked tree to what a published site is actually made of, before anything is pointed at it:

unzip -q "$tmp/pages.zip" -d "$tmp/out"
find "$tmp/out" ! -type f ! -type d -delete
chmod -R a-s,go-w "$tmp/out"

Info-ZIP already refuses the two obvious escapes — it strips ../, it warns stripped absolute path spec from /x, and it will not write through a symlink. What it does do is restore one, target and all, and the cp -a that follows preserves it. That is enough on its own, because caddy’s file_server follows a symlink out of its root, and modules/acme.nix group-owns the certificate directory by caddy so that caddy can read key.pem. One ln -s /etc/caddy/certs/key.pem x in a published artifact would otherwise serve the apex certificate’s private key — and every one of its eleven SANs — at https://pages.<domain>/<owner>/<repo>/x.

A published site is files and directories; a symlink in one has no legitimate use here. They are deleted rather than rejected so that one malformed artifact cannot wedge a repo’s publishing, and the match is on the entry itself and never its target, so this cannot follow a link out of $tmp.

The filter is an allow-list of entry types, not a list of known-bad ones. A zip cannot carry a device node or a socket, but it can carry a fifo, and cp -a would preserve it into the served tree — where caddy’s file_server blocks forever on open(2) the first time anyone requests that path. Deleting everything that is not a regular file or a directory also means the next entry type an archive format learns does not get a free pass.

The chmod covers the other thing a zip carries: a mode. Info-ZIP already drops setuid and setgid on extraction, so a-s is belt-and-braces against an extractor that someday does not — but go-w is not redundant. A world-writable file in the pages volume is one that any process on this host can rewrite after publication, with caddy serving whatever it finds on the next request. It runs after the delete, so there is no symlink left for chmod -R to follow out of $tmp.

checks.pages-pull-strips-symlinks holds both: it asserts the strip sits between the unzip and the copy and that the chmod sits between the strip and the copy — a chmod ordered before the strip would walk a tree that still contains symlinks and follow one out, reintroducing the exact escape the strip exists to close. It then evals those exact lines — lifted out of the evaluated unit, not retyped — against a tree unpacked from a hostile zip, with a fifo and a world-writable file planted on top of it.

Why the unit is still root

pages-pull runs as root, and the reason is the destination rather than the work. The tree lands in /var/lib/docker/volumes/<vol>/_data, and /var/lib/docker is 0710 root:root — nothing unprivileged can even traverse it. Adding the unit’s user to the docker group is not the alternative: that group is root-equivalent by design. Dropping the privilege properly means moving pages off a docker volume onto a plain bind-mounted directory, which is a volume migration rather than an edit.

So root is bounded instead. The unit parses a network-fetched ZIP with unzip, which is the one place in this module where hostile content meets a parser, and the mitigation that matters is that winning there reaches nothing worth having. Under ProtectSystem=strict with ReadWritePaths naming only /var/lib/docker/volumes, plus PrivateTmp, NoNewPrivileges, an empty CapabilityBoundingSet and SystemCallFilter=@system-service, a compromised unzip can write to the pages volume and its own private /tmp and nothing else — not /var/lib/sops-nix, not /etc, not /home, not the docker socket, with no way to gain a capability back.

ReadWritePaths names the volumes parent, not the volume’s own _data path, on purpose: _data does not exist until docker first creates the volume, and a ReadWritePaths entry that does not resolve fails the unit on a fresh box.

When it runs

A Forgejo system webhook on action_run_success, so publishing is an event rather than a poll. The whole path is:

triggerone system hook in Site Administration, firing for every repo on the instance
targethttp://<dockerBridgeGateway>:<pagesHookPort>/hooks/pages-pull
authHMAC-SHA256 over the body, read from X-Hub-Signature-256
receivermodules/pages-hook.nix — webhook(1) in a container on the proxy network, no published port
effecttouches one file; a systemd .path unit starts pages-pull as root

It never leaves the proxy network. Both Forgejo and the receiver are containers on it, so the delivery is container-to-container. There is no caddy site, no published port, no public listener — and no firewall rule at all, because the port never exists on the host.

That is also the correction to an earlier claim here: this used to say a webhook cost “an HTTP receiver on the mail server, which is not a trade worth making”. It assumed the receiver had to be public.

A system hook, not a per-repo hook. Per-repo would reintroduce exactly the per-repo setup step that discovery removed. One hook covers everything, including repos that do not exist yet.

The receiver parses nothing. Any successful Action Run pokes pages-pull, which is idempotent and cheap. Reading the payload would trade a slightly smaller number of no-op runs for a coupling to Forgejo’s ActionPayload schema.

The listener holds no privilege. It runs as its own uid from modules/ids.nix, with a read-only rootfs and every capability dropped, and its entire capability is touching one file in a bind-mounted directory; the .path unit does the privileged half. A network-facing process that can run systemctl start is a network-facing process that is root-adjacent.

The hourly timer stays, and is now a safety net rather than the mechanism. A webhook is a delivery and deliveries are lost — the receiver can be down mid-deploy, Forgejo’s retries can run out, the hook can be switched off in a web form nothing here can see. Each of those leaves a site frozen with no error anywhere. The sweep makes the worst case “stale for up to an hour”.

Forgejo has to be told the destination is allowed

ALLOWED_HOST_LIST defaults to external, which permits public addresses and blocks private ones — so out of the box Forgejo refuses to deliver and the webhook silently never fires. modules/containers/forgejo.nix sets it to pages-hook.

The container name, not an address, and that is the security-relevant part. The list matches hosts, not host:port. While the receiver ran on the host, this had to name 172.17.0.1 — which also permitted a webhook aimed at anything else bound there, and grafana (3000), syncthing’s GUI (8384), tempo (4317/4318) and pgbouncer (6432) all bind 0.0.0.0 and answer on it. Forgejo webhooks can use GET and record the response body in their delivery history, so that was a read primitive with an exfiltration channel: whoever could create a webhook could read tailnet-only services without being on the tailnet.

A name works because the matcher is MatchHostName(host) || MatchIPAddr(ip) — a name pattern alone is sufficient, so no private address needs allowing and 172.17.0.1 stops matching at all. One destination, nothing else.

The failure mode is worth knowing because everything on the RECEIVING side looks correct while it happens: the socket is listening, the rule matches, and there is no dropped packet, no connection refused and nothing in the receiver’s journal — because nothing is ever sent.

Look at the sender instead. Forgejo logs the refusal as an error on its own service:

journalctl -u docker-forgejo | grep -i 'unable to deliver webhook'

and the hook’s settings page shows the same thing under Recent deliveries.

The one hand-kept value

The hook’s Target URL and secret live in a web form, so nothing in this repo can verify they match infra.pagesHookPort and forgejo/system_webhooks/pages_pull/secret. A mismatch is at least loud in two places: a 403 in journalctl -u pages-hook, and a failed delivery in the hook’s own history in Site Administration.

Failure isolation

Each repo runs in its own subshell, so one bad repo — renamed, deleted, made private, a network blip, a corrupt zip — cannot take down the rest of the loop. A transient failure heals itself on the next tick; a persistent one is the failure worth guarding against, because left unguarded it would permanently block every repo listed after it, and silently: the unit “succeeding” on the repos before the broken one looks no different from everything being fine.

So the loop logs it, keeps going, and fails the unit at the end. failed is set, not incremented — whether it was one repo or all of them, the answer is “look at the journal”, not a count.

journalctl -u pages-pull --since -1h
systemctl list-timers pages-pull

Checking it worked

curl -sSI https://pages.hu-tao.dev/hutao/vps/docs/ | head -1

Runbook

Commands, in the order you are likely to want them. Everything here runs on the VPS over ssh on port 2222 (ssh -p 2222 hutao@vps); the runner’s equivalents are at the bottom, and they are reached differently.

Is anything broken

systemctl --failed
systemctl list-units 'docker-*'
journalctl -u docker-forgejo -f           # every container logs to the journal

Everything else

systemctl list-units 'docker-*'
journalctl -u docker-forgejo -f           # every container logs to the journal

systemctl status acme-renew-hu-tao.dev.timer  # renews well before expiry; a no-op most days
systemctl start acme-hu-tao.dev.service   # force a renewal check

systemctl status restic-backups-b2.timer restic-backups-minecraft.timer
restic-b2 snapshots                       # wrapper with the repo and password wired in

systemctl status postgresqlBackup.timer   # 23:15, deliberately BEFORE restic's 00:00-01:00 window
systemctl start postgresqlBackup          # dump every database now
psql -h 127.0.0.1 -p 6432 -U serenity serenity_bot   # through pgbouncer, from the tailnet

fail2ban-client status forgejo-ssh
nft list table inet f2b-table             # where the bans actually are
nft list table inet nixos-fw

tailscale status                          # peers, and this node's own address
tailscale whois 100.109.115.12            # THIS node: `Tags: tag:vps` or the ACL does not apply

systemctl status image-archive.timer      # 23:30, ahead of restic's 00:00-01:00 window
systemctl start image-archive             # docker save every pullable image now
ls -la /var/lib/image-archive             # one .tar + one .id per image

When an image cannot be pulled

Nothing to do — every pullable container’s unit runs ensure-image before it starts, which tries the local image, then a pull, then the nightly archive, then restic. journalctl -u docker-<name> names whichever step it reached.

The recovery path can be exercised by hand without touching a running service:

restic-b2 restore latest --path /var/lib/image-archive \
  --include /var/lib/image-archive/<slug>.tar --target /tmp/rt

--path is not optional: this repository also holds the weekly minecraft snapshots, and a bare latest can name one of those, which carries no archive. --target / is what ensure-image uses, because restic recreates the absolute path under the target. The slug is the image reference with /, : and @ each replaced by _. Verified end to end on 2026-09-21.

The weekly minecraft job stops both worlds’ servers, snapshots, and starts them again from ExecStopPost — so they come back whether restic succeeded or not. A live world is not consistent on disk: the server holds region files open and writes them in place, which is why the daily backup excludes that volume and this job exists.

On the runner

The runner is off the tailnet on purpose, so every command goes through the VPS:

ssh -J hutao@vps:2222 root@46.225.61.172

systemctl status forgejo-runner
journalctl -u forgejo-runner -f     # a job that never starts shows here
journalctl -u forgejo-runner-identity     # the uuid/secret compose step

podman ps                                 # job containers, one per running job
podman images                             # localhost/forgejo-ci-nix-node must be here
systemctl restart forgejo-runner-ci-image # reload it if the prune ate it
du -sh /var/lib/forgejo-runner/cache      # the Actions cache
nft list table inet nixos-fw              # the one-way rules; counters included

nft list counters is the quick check that the one-way rule is doing something: vps_allowed_out should climb while jobs run, and vps_blocked_out should stay where it was. See CI runner isolation.

The ingress pair answers the other direction — ssh_from_vps climbs every time you open the jump above, and ssh_blocked is anyone else trying:

ssh -J hutao@vps:2222 root@46.225.61.172 nft list counter inet nixos-fw ssh_from_vps
ssh -J hutao@vps:2222 root@46.225.61.172 nft list counter inet nixos-fw ssh_blocked

Renovate

Renovate runs on the VPS, not the CI runner: its token can write across hutao/* and skavex/*, and a runner treated as hostile never holds it.

systemctl status renovate.timer           # daily 12:00 UTC, ±15 min
systemctl start renovate                  # run now
journalctl -u renovate -n 200

Nothing opens until it is ticked on the Dependency Dashboard (dependencyDashboardApproval in renovate.json5). With the CVE scanner dropped, reading that issue is the only path for security updates.

Failure modes and recovery

The design is shaped by which failures roll back automatically and which do not. This page is the table of both, and then the three that are recent scars.

FailureCaught byRecovery
firewall / sshd / networking change locks you outdeploy-rs auto-rollbackautomatic
unbootable kernel / initrdGRUB generation menu (5 s at boot)pick the previous generation
a container fails to startnothing — deploy confirms reachability, not healthnixos-rebuild --rollback or fix forward
nftables reload wiped docker’s chainsnothing automatic; the symptom is the next container start failingsystemctl restart docker, and keep flushRuleset = false
a rotated sops secret didn’t reach a containernothing — the container holds a stale inodecopy-to-stable-path or env-file, see Secrets
tofu plan shows an unexpected diff on unchanged infrayou, reading the planthe state is wrong, not the infra — never apply; refresh/import, verify on the box
a fail2ban jail bans a docker/bridge addressnothing — a forward-chain reject on an internal IP downs every containersystemctl stop fail2ban; keep private ranges in ignoreIP
a deploy restarts every container at once and one racy unit exits non-zerodeploy-rs aborts — and its deactivation stops every containersee the abort that was worse than the failure
tofu apply fails with test(s) failed (400) on the tailnet policythe policy’s own tests, run by Tailscale at apply timethe narrowing is not live yet — see the bootstrap deadlock
a container image is gone from the registrynothing — the pull fails at the next start, on an image that has worked for monthsautomatic: ensure-image falls back to the nightly archive, then to restic — see the image archive
the DNSSEC chain breaksnothing here — kuma-check probes from the VPS, whose resolver may not validatedig +dnssec @1.1.1.1 hu-tao.dev SOA from off-box; until then the domain is dark for validating resolvers

The fail2ban blast radius

A forward-chain fail2ban jail has a blast radius the size of the whole box. If it ever bans an internal source it rejects all forwarded traffic, not one attacker — every container goes dark at once.

That is why the sites are rate-limited in caddy rather than banned in nftables: an in-process limiter can only throttle, it cannot take the forward plane down. See Access control.

The host’s ignoreIP includes 172.16.0.0/12, which holds every docker network here, so no jail can ban a bridge address at all — structurally, not because a regex happens not to match one. docker-mailserver’s own fail2ban carries the same exemption, for the same reason inside its own namespace.

Before adding any forward-chain jail, ask what happens when it bans 172.30.0.x.

The abort that was worse than the failure

The least intuitive row, and the widest. A failed activation does not leave the box on the previous generation running. It leaves it on the previous generation’s configuration, with nothing started.

flowchart TB
    u["nix flake update<br/>new nixpkgs"] --> r["every unit changes<br/>→ every container restarts at once"]
    r --> race["one racy unit exits non-zero<br/>(3-second transient)"]
    race --> abort["deploy-rs samples unit state,<br/>sees one failed unit, ABORTS"]
    abort --> deact["deactivation stops<br/><b>every</b> container"]
    deact --> out["15 healthy services down, mail included<br/>6 minutes"]
    deact --> lock["switch-to-configuration refuses:<br/>Could not acquire lock"]

    classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
    class abort,deact,out,lock bad

On 2026-09-16 a three-second transient in one non-critical container therefore cost fifteen healthy services, mail included, for six minutes. The abort was more destructive than the failure it was responding to.

Recovery, in order:

systemctl restart docker
# then start the container units by hand — switch-to-configuration will refuse
# with "Could not acquire lock" while deploy-rs still holds it

The structural fix is to order fragile units behind a readiness gate rather than letting them race the restart. Full writeup: 2026-09-16 — flake update rollback.

tofu state recovery

State is gitignored — it holds every value tofu has ever read. A clone has none, and apply from no state builds a second server and moves DNS to it.

tofu/imports.tf prevents that declaratively: an import block per live resource, inert while state tracks it, active when it does not. tofu init && tofu plan then rebuilds state.

A correct recovery plan reads 32 to import, 0 to add, 1 to change, 0 to destroy — the one change being three provider-side booleans on hcloud_server.vps that the importer never sets. Anything else means stop.

Expect a second change, tailscale_acl.main, until the policy in this repo has been applied at least once: the import reads whatever the tailnet is currently serving, and the diff against tailscale-policy.hujson is the point. Read it before applying — that resource replaces the whole document, so anything the live policy carries and the file does not is deleted with nothing in the plan naming it as a loss. See The tailnet policy.

Three resources are in state that nothing declares — hcloud_network, hcloud_network_subnet and hcloud_server_network.vps, left over from the private network that was deleted. They cannot be given import blocks (Configuration for import target does not exist), and plans refresh them and then omit them from resource_changes entirely, which is recorded in tofu/imports.tf as not understood rather than explained away. Clearing them means tofu state rm.

The sharp edge (learned the hard way, 2026-09-05): the hcloud provider never reads public_net into state, so a post-import plan proposes adding it — and on this resource that detaches the primary IPs before reattaching. Applying it once took the mail IP off a running host. server.tf now carries lifecycle.ignore_changes = [public_net], so the block can never become an action. The primary IP is its own resource with delete protection precisely so a mistake here is minutes of downtime rather than a lost address.

A plan that disagrees with what you know to be true is a state problem, not an infrastructure problem. Never reach for apply to make it agree.

Runner-specific failures

SymptomCauseFix
jobs queue and never startthe runner cannot reach Forgejo, or its identity is wrongssh -J hutao@vps:2222 root@<runner>, then journalctl -u forgejo-runner
403 not permitted from a CI runner in a jobthe workflow hit a path outside the four-path allow-listwiden the allow-list deliberately, or change the workflow — see runner isolation
upload-artifact fails with no HTTP request at allGitHub’s action, not the Forgejo forkforgejo/upload-artifact@v5
you cannot ssh to the runnerits route in is defined by the VPS’s outbound rulesdeploy the VPS first; the runner has no tailnet and no other path
every job on nix-node fails in ~3 s, on unchanged workflowsthe daily podman system prune --all deleted the locally built job imagesystemctl restart forgejo-runner-ci-image — or just deploy; both now reload it. See the garbage collector eats it

A runner that is broken past recovery is replaced, not repaired — it holds a nix store and a job cache and nothing else. Rebuild in place, never destroy and recreate: CX instance types are limited availability.

Development

nix develop            # or `nix develop -c zsh`

The shell carries everything: nixfmt, statix, sops, age, nixos-anywhere, nixos-rebuild, opentofu, deploy-rs, pre-commit, gitleaks, markdownlint-cli2 and mdbook. .#ci is a deliberately smaller subset — see CI.

Running it locally

Tests:

nix run github:nix-community/nixos-anywhere -- --flake .#vps --vm-test

The actual system (infinitely more useful):

QEMU_OPTS="-vnc :0" nix run .#default

And ssh into it from another terminal — ssh -p 2222 hutao@127.0.0.1 (there’s no place like 127.0.0.1). Port 2222 on both ends, because sshd moved off 22 so forgejo could publish it.

The VM disk is 32G (virtualisation.vmVariantWithDisko). The disko default of 2G leaves ~987M for / once the ESP takes its gigabyte, which cannot hold the sixteen declared images — the VM fills up mid-boot and every service that needs disk fails in a way that reads like a bug in that service. This applies to nix run . only; the Hetzner disk is sized by the provider.

Working from a Mac

The devShell is built for all four mainstream systems — x86_64-linux, aarch64-linux, aarch64-darwin, x86_64-darwin. Every tool in it, nixos-anywhere, nixos-rebuild and deploy-rs included, exists on each. Secrets, formatting, statix and the hooks work unchanged.

What does not carry over is building the system closure. The outputs that describe the box — nixosConfigurations, packages, apps — are x86_64-linux only, so nix build ., nix run . (the QEMU VM) and nix run .#install fail on anything else: a Mac has no Linux builder at all, and an aarch64-linux workstation is the wrong architecture. Deploys work, but only if the build happens somewhere else:

deploy -s --remote-build .#vps         # build on the VPS itself
nixos-rebuild switch --flake .#vps-hetzner \
  --target-host hutao@vps --build-host hutao@vps --use-remote-sudo

-s is not optional here. deploy-rs runs nix flake check first, and every check in this flake reaches nixosConfigurations.*.system.build.toplevel, so the check itself is an x86_64-linux build:

error: build of '…-10-acme.conf.drv^*' failed: platform mismatch
       Required system: 'x86_64-linux'   Current system: 'aarch64-darwin'

There is nothing to keep by skipping selectively — checks.aarch64-darwin exists but both entries depend on the same Linux closure, so none of them build here either.

--remote-build then evaluates locally (which darwin does fine), copies the .drv with nix copy --to ssh-ng://hutao@vps --derivation, and realises it on the box. That copy needs the ssh user to be a trusted nix user; hutao is in wheel and modules/nix.nix trusts @wheel, so it already is.

The alternative is a Linux remote builder in /etc/nix/machines (or nix-darwin’s nix.linux-builder), after which the plain commands above work as written — including nix run .#install, which is otherwise Linux-only and so still the reason a first install is done from a Linux machine.

Formatting

Formatting is nixfmt, not nixpkgs-fmt — every .nix file here conforms to it and the two disagree on multi-argument lambdas, so the wrong one reformats the whole tree.

nixfmt $(git ls-files '*.nix') && statix check .
tofu -chdir=tofu fmt -recursive && tofu -chdir=tofu validate

Hooks

nix develop -c pre-commit install         # once per clone
nix develop -c pre-commit run --all-files

.pre-commit-config.yaml is the single definition — the local commit hook and CI run the same file, so a check cannot pass here and fail there. It covers nixfmt, statix, tofu fmt, markdownlint-cli2, the usual whitespace/YAML/merge-conflict hooks, and gitleaks over the staged diff.

See below for what the markdown hook does and does not do.

gitleaks scans the staged diff rather than the working directory on purpose: gitleaks dir reads gitignored files, and tofu/terraform.tfvars legitimately holds live tokens — scanning it would fail the hook forever over a file git will never accept. History scanning is a CI step instead.

Writing documentation

This book is docs/. mdbook and mdbook-mermaid are in the dev shell:

nix run .#docs                             # install assets, then serve
nix develop -c mdbook build docs           # what CI does

nix run .#docs exists because the build has a prerequisite that is easy to forget: mdbook-mermaid install docs writes mermaid.min.js and mermaid-init.js next to book.toml, and book.toml references them. Those two files are gitignored — 2.6 MB of vendored minified JS whose version is already pinned by flake.lock — so a fresh clone does not have them and mdbook build fails until they are written. The app does both steps; the workflow does the same two commands explicitly.

create-missing = false in book.toml, so a link to a page that does not exist fails the build rather than publishing a 404.

Diagrams are mermaid, and that is the point

Every diagram here is a ```mermaid fence in the markdown. Forgejo bundles mermaid 11.16.1, so the same source renders in three places with no build step: this book, the Forgejo web UI when browsing docs/src/, and a pull request that changes one.

That is the whole reason they are not SVGs, D2, or the ASCII art they replaced. A picture that only exists after a build is a picture nobody sees while reviewing the change that invalidates it.

Markdown is linted, not formatted, by the hook

markdownlint-cli2 reports a code fence with no language and a paragraph past 80 columns. It rewrites nothing.

Formatting is prettier’s, run by an editor on save rather than by a hook, and .markdownlint-cli2.yaml is tuned so prettier’s output passes untouched. The config lists every rule that is off and why. Prettier is deliberately not in the dev shell: nixpkgs-26.05 carries 3.8.3, which mangles a paragraph containing both an intraword underscore and an emphasis span.

docs/superpowers/ is not linted — those are agent-generated plan and spec records, kept as history.

Checking the Minecraft mod lists

nix run .#minecraft-mod-check

It checks every declared Modrinth slug against the exact Minecraft version and loader each world runs, read from the evaluated config rather than a copy of it. CI builds the same thing as .#minecraft-mods, which needs --option sandbox false because it uses the network.

CX33 → CX33 (→CX43) migration

Status: completed. This is the plan as it was run, kept as the record. The old box (137766340) has since been deleted, so the rollback in step 5 no longer exists and legacy_server_ids in tofu/variables.tf is empty. The serenity bot, which this plan left on the old box, was ported afterwards in 8bba9bb. For a future move, reuse the method (the data mapping, populate volumes before first start, the primary-IP handover), not the numbers.

Moving the stack from ubuntu-4gb-fsn1-2 (server 137766340, Ubuntu + Ansible) to ubuntu-8gb-fsn1-1 (server 163906050, NixOS from this flake). Both in fsn1, which is the fact the whole plan rests on: primary IPs are location-bound, so 167.233.24.58 and its smtp.hu-tao.dev PTR can move between these two machines. Nothing in DNS changes, SPF keeps naming the same address, and sending reputation carries over intact.

Rescale to CX43 came after the migration. The box is CX43 today.

Data mapping

The old box bind-mounts host directories; this config uses named docker volumes. That makes the transfer a mapping, not an rsync of one tree. ~27 GB total.

Source (old box)SizeDestination (new box)Entries
/srv/minecraft/data/15 Gvolume minecraft_data8107
~/syncthing/11 G~/syncthing/ (host path)—
/srv/forgejo/data/271 Mvolume forgejo_data1151
~/data/navidrome/123 Mvolume navidrome_data5020
/srv/mailserver/data/dms/mail-state/95 Mvolume dms_state—
volume grafana-data50 Mvolume grafana_data—
volume tempo-data16 Mvolume tempo_data—
/srv/kuma/data/8.7 Mvolume kuma_data9
/srv/mailserver/data/dms/mail-data/2.5 Mvolume dms_mail—
/srv/mailserver/data/roundcube/db/1.3 Mvolume roundcube_db—
/srv/mailserver/data/dms/config/12 Kvolume dms_config1 account
~/.config/syncthing/small~/.config/syncthing/ (host path)—

Read these sizes as root. Measured as hutao, /srv/mailserver/data reports 15 M; as root it is 110 M, because 24 paths under /srv are owned by container uids (the mail store is uid 5000) and find/du silently skip what they cannot read. A copy run as hutao therefore loses mail without erroring. Root over Tailscale SSH is what makes the copy complete — ssh root@<tailnet-ip> works even though sudo on that box demands a password.

mail-state is 95 M and is the bulk of the mail data. It is not scratch: it holds dovecot’s indexes and UIDVALIDITY, and rspamd’s trained bayes database. Skip it and every IMAP client re-downloads everything and your spam filter starts from zero.

Deliberately not transferred:

SkippedWhy
/srv/certbot/datasecurity.acme issues a fresh certificate; the certbot lineage layout is not the same and is not read any more
/srv/dozzle/datajust users.yml, rendered from sops by dozzle-users.service
/srv/caddy, caddy_caddy_* volumesACME state caddy no longer manages
/srv/mailserver/data/dms/mail-logs11 M of logs
roundcube/configrendered from the Nix store
postgres-data, serenity-discord-bot_postgres-databoth 0 bytes, 0 links — dead volumes
deploy_* volumes, /srv/camofox-browserexpenses app and camofox, staying on the old box
docker build cache3.8 G of nothing

~/.config/syncthing carries the node’s TLS keypair, which is its device ID. Copy it and every paired device keeps working; regenerate it and you re-pair by hand on every phone and laptop.

The Discord bot’s data

Not in a docker volume, which is why the two postgres volumes on the old box are 0 bytes: the bot talks to the host’s PostgreSQL 18.6, listening on 127.0.0.1:5432 and on the tailnet address. DATABASE_URL in the bot’s .env points at database serenity_discord_bot, 8.4 MB, 9 tables — and only user_stats has rows (258 of them). Everything else is empty schema.

Dumped read-only with pg_dump --no-owner --no-privileges, and held in two places so it does not live only on the box being retired:

old box:     ~/migration-dumps/serenity-<timestamp>.sql
workstation: ~/migration-dumps/serenity-<timestamp>.sql

The bot stayed on the old box through the cutover and was ported afterwards in 8bba9bb: host services.postgresql behind pgbouncer, the dump restored, and its credentials in sops. --no-owner --no-privileges is what let the restore land under a different role than the Ubuntu one that owned it.

The two Cloudflare tokens

There are two, with different jobs, and confusing them costs an hour:

WhereUsed byToken ID
secrets.yaml → cloudflare.api_tokenlego / security.acme on the VPSa13f8e29ba9ebe4201e5aef3d1723ec7
tofu/terraform.tfvars → cloudflare_api_tokentofu, from a workstationc91db61e8b430d39e780fd5e6098c225

The VPS token carries an IP filter pinned to the old server, so DNS-01 fails from anywhere else with 403 9109: Cannot use the access token from location: <ip>. security.acme then falls back to a self-signed certificate and starts dependent services anyway, so caddy and DMS come up serving a placeholder and nothing looks broken until you check the issuer:

ssh -p 2222 hutao@<host> sudo cat /var/lib/acme/hu-tao.dev/cert.pem \
  | openssl x509 -noout -issuer
# issuer=CN=minica root ca …   <- placeholder
# issuer=C=US, O=Let's Encrypt …  <- real

Test the token directly rather than by triggering ACME — Let’s Encrypt caps failed validations at 5 per account per hostname per hour, and lego burns one per attempt:

curl -sS -H "Authorization: Bearer $TOKEN" \
  'https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev'

This resolves itself at cutover, when the box inherits the old server’s address.

Order of operations

Volumes must exist and be populated before their container first starts, otherwise docker creates them empty and the service initialises itself blank — forgejo would make a new instance, DMS an empty mail store. So: install, stop the containers, load the data, start.

Observed on the freshly installed box, and both are the CORRECT empty-state behaviour rather than faults to chase:

  • forgejo serves HTTP 200 but /api/v1/version 404s and repos: 0 — it is a blank instance that has not been through its install wizard. Restoring forgejo_data is what makes it the real instance; do NOT click through the wizard first, or you create a second one.
  • DMS loops on You need at least one mail account to start Dovecot (120s left...) and then exits, so systemd restarts it. Accounts live in postfix-accounts.cf inside the dms_config volume. Until that volume is restored there is no account, and DMS refuses to run Dovecot or Postfix — which is why ports 25/465/587/993 have no banner yet. Restoring dms_config resolves it; nothing needs fixing.
export SSH_AUTH_SOCK=/tmp/hutao-agent.sock   # key is passphrase-protected
NEW=178.105.223.159

1. Install NixOS

nix run github:nix-community/nixos-anywhere -- \
  --flake .#vps-hetzner \
  --target-host root@$NEW \
  --extra-files <staged age key dir>

vps-hetzner, not vps: it targets /dev/sda. The extra-files directory carries /var/lib/sops-nix/key.txt at 0600 — without it the host boots unable to decrypt anything, including its own root password.

2. Quiesce, then load data

ssh -p 2222 hutao@$NEW 'sudo systemctl stop "docker-*"'

Old box first, so nothing is written mid-copy:

ssh hutao@100.97.90.108 'sudo systemctl stop docker-services; \
  cd /srv && for d in forgejo mailserver minecraft kuma navidrome; do \
    (cd $d 2>/dev/null && sudo docker compose down); done'

Then, from the OLD box over the tailnet (traffic stays inside fsn1 rather than going via a workstation), for each row of the mapping:

sudo rsync -aHAX --numeric-ids --info=progress2 \
  /srv/forgejo/data/ root@hu-tao:/var/lib/docker/volumes/forgejo_data/_data/

-aHAX --numeric-ids because uid/gid must survive verbatim: forgejo’s repos, the mail store and grafana’s data are owned by container uids that mean nothing in either host’s /etc/passwd.

Create each volume before writing into it:

ssh -p 2222 hutao@$NEW 'for v in forgejo_data dms_mail dms_state dms_config \
  roundcube_db minecraft_data kuma_data navidrome_data grafana_data tempo_data; \
  do sudo docker volume create $v; done'

3. Verify before cutover

DNS-01 works regardless of which IP the box holds, so certificates can be issued and the whole stack tested while the old box is still live.

ssh -p 2222 hutao@$NEW 'systemctl start docker-network-proxy; sudo systemctl start "docker-*"'
ssh -p 2222 hutao@$NEW 'systemctl is-system-running; systemctl --failed'
curl --resolve git.hu-tao.dev:443:$NEW https://git.hu-tao.dev/api/v1/version
curl --resolve status.hu-tao.dev:443:$NEW https://status.hu-tao.dev/api/entry-page

Check journalctl -u docker-mailserver shows the certificate loading from /certs, and that forgejo lists the migrated repositories.

4. Cutover — the IP handover

Hetzner requires a server to be powered off to (un)assign a primary IP.

  1. Final delta rsync of the mapping rows (minutes, since only deltas move)
  2. Power off both servers
  3. Unassign 147045244 (178.105.223.159) from 163906050
  4. Unassign 134632948 (167.233.24.58) from 137766340
  5. Assign 134632948 to 163906050
  6. Assign 147045244 to 137766340 — the old box stays online at the other address, so serenity-bot, camofox and the expenses app keep running
  7. Power on both

smtp.hu-tao.dev follows 134632948 automatically; it is a property of the IP.

5. Rollback

Steps 2–7 in reverse. The old box is untouched, still holds its data, and has delete and rebuild protection on. Rollback is ~5 minutes and costs nothing but the swap.

(Historical: 137766340 has since been deleted, so this rollback no longer exists.)

After it settled

All done. Kept as the checklist that was run:

  • ✅ Attach firewall 11483636 to 163906050
  • ✅ Enable delete + rebuild protection on 163906050 (tofu/server.tf)
  • ✅ Import into tofu: server, primary IPs, firewall, rDNS, and the Cloudflare records (tofu/imports.tf). Read the plan: abort if it shows destroy and then create on hcloud_server
  • ✅ Empty legacy_server_ids in tofu/variables.tf once the old box is retired
  • ✅ Port serenity-bot (8bba9bb, modules/containers/serenity-bot.nix)
  • ✅ From then on deploys are deploy .#vps, with auto-rollback on lockout

What changed in the port

Not a 1:1 translation. The deliberate departures:

  • Data lives in named docker volumes, never a bind-mounted host directory. services.restic backs up /var/lib/docker/volumes wholesale, so a service added later is covered the moment it declares a volume. A backup that has to be told about each new service is a backup that eventually stops covering one. The Ansible roles’ data guards (assert that data/world exists before provisioning) have no equivalent and need none — there is no path to point at the wrong place.

  • Configuration comes from the Nix store, read-only. Store files are 0444, which is what the tempo (uid 10001) and grafana (uid 472) permission failures in the Ansible setup were about; and a store path changes when its content does, so systemd recreates the container on a config-only change. That is what recreate: always was working around.

  • Certificates are security.acme, not a certbot container plus cron. The certificate is named hu-tao.dev with the subdomains as SANs, so reordering the list cannot silently issue a second lineage the way certbot’s name-after-the-first--d behaviour could. Note NixOS calls the key key.pem, not certbot’s privkey.pem, and DMS reads it with SSL_TYPE=manual rather than guessing from $SSL_DOMAIN.

  • Images pin a release tag and no digest. The Ansible repo pinned tag@sha256:…; the digests have since gone stale, and a release tag has been the more stable of the two in practice. :latest is still never used — for tempo it is a main-branch build that reports a version which was never released.

  • Real credentials everywhere. Grafana’s anonymous-Admin access is off and the login comes from sops, as do dozzle’s bcrypt hash and minecraft’s RCON password. uptime-kuma is the exception, unavoidably: it has no environment variable or config file that seeds an admin account — the first visitor is prompted to create one and the route then closes. Create it immediately after the first deploy. searxng is the other exception, for the opposite reason: it has no concept of a user at all, so there is nothing to seed — caddy’s basic_auth is the entire access control and the hash lives in sops. Because basic_auth runs a cost-14 bcrypt on every request, caddy’s rate_limit (a compiled-in module) sits in front of it and returns 429 before the hash runs, so a password flood cannot become CPU exhaustion. It is in-process on purpose — see the fail2ban row in Failure modes for the forward-chain ban whose blast radius it avoids.

    Generate that hash with mkpasswd, which is already on the host:

    mkpasswd -m bcrypt -R 14
    

    -R 14 is not optional. mkpasswd defaults to cost 05 and caddy’s own hash-password uses 14, so the default silently produces a hash 512x cheaper to attack than the one caddy would have made. The $2b$ prefix mkpasswd emits is fine — caddy verifies through golang.org/x/crypto/bcrypt, which records the minor version without validating it, so $2a$, $2b$ and $2y$ are interchangeable.

  • Not ported: camofox. It is a host systemd unit for an npm project checked out under /home, not a container service, and it depends on a tree this image does not create. serenity-bot was in the same position and was ported later as a container built from upstream’s Dockerfile; see modules/containers/serenity-bot.nix.

Postmortems

Written when something broke badly enough that the fix is not obvious from the diff, and kept afterwards because the reasoning is the part that does not survive in git.

The house style: a timeline with real timestamps, impact measured rather than estimated, and a “what went badly” section that is allowed to be unflattering. A postmortem that only lists what went well is a press release.

DateTitleOne-line cause
2026-09-16A dependency window that took the box down for six minutesa three-second transient in one non-critical container made deploy-rs abort, and the abort stopped every container

The structural lesson from that one is in Failure modes: a failed activation leaves the box on the previous generation’s configuration, not its previous running state.

2026-09-16 — a dependency window that took the box down for six minutes

Impact: ~5m44s, 22:19:55–22:25:39 EEST. Every service except the two minecraft worlds was down, mail included. Trigger: a nix flake update deploy. Root cause: a known, documented, deferred missing ordering dependency on docker-forgejo-runner.service. Detected by: kuma-check, correctly, within 13 seconds of its next tick.


1. What we set out to do

Clear all seven pending items on the Renovate Dependency Dashboard (#12) inside one 30-minute maintenance window: three container bumps, one container major (docker-mailserver 15.1.0 → 16.0.1), lockFileMaintenance, and two tofu provider bumps.

Six of the seven landed. The seventh took the box down on its way in.

2. Timeline

All times EEST. The box logs UTC; subtract three hours.

TimeEvent
22:00:40Cold-stop grafana + navidrome. Local tar, then restic a3534fdf.
22:02:44Deploy 1 (gen 58): dozzle v11.1.0, grafana 13.2.2, navidrome 0.64.0. Clean, 40s.
22:11:24Cold-stop mail. Local tar, then restic 63564755.
22:12:43Deploy 2 (gen 59): docker-mailserver 16.0.1 + the opendkim gid fix. Clean, 37s.
22:16:39Deploy 3 starts: flake.lock only.
22:19:28New nixpkgs changes every unit → every container restarts at once. Runner stopped.
22:19:52Runner starts, declares against https://git.hu-tao.dev/, caddy is not listening yet: fail to invoke Declare … connection refused. Exits 1.
22:19:53deploy-rs samples unit state, sees one failed unit.
22:19:55Deploy 3 aborts. De-activation stops every container. Outage begins.
22:19:58The runner’s own Restart=always brings it up successfully — 3 seconds after the abort decision.
22:20:08kuma-check fails: Failed to connect to status.hu-tao.dev:443.
22:25:01kuma-check fails again.
22:25:03systemctl restart docker (the documented recovery).
22:25:09switch-to-configuration switch → Could not acquire lock. Fell back to starting units directly.
22:25:39Mail answering. Outage ends.
22:30:27kuma-check green again.
22:33:16Deploy 4 (gen 60): flake.lock + the readiness gate. Clean, 40s.
22:34:54Reboot complete on kernel 6.18.52. 38 seconds down.
~22:38tofu apply: 0 added, 0 changed, 0 destroyed.

3. Impact, measured

  • Mail (25/465/587/993) refused connections for 5m44s. No bounces or deferrals are attributable to the window. That is expected but only weak evidence: while the container was down we logged nothing, because nothing reached us. Sending MTAs queue and retry on schedules measured in hours, and six minutes is far inside every one of them, so loss is very unlikely — but it is unprovable from our side.
  • Everything behind caddy — forgejo, webmail, kuma, navidrome, searxng, pages — refused connections for the same period.
  • Observability (grafana, tempo, dozzle) down, and kuma down, so the public status page could not report its own outage. This is the paradox Observability anticipates, and the mitigation worked: see §6.
  • minecraft / minecraft2 restarted and came back on their own. The other fourteen containers did not.

4. Root cause

docker-forgejo-runner.service has a hard runtime dependency on caddy that is nowhere expressed to systemd.

The runner’s first action on start is declaring itself against instanceUrl, which is the public https://git.<domain>/ — deliberately, because the URL is handed to job containers that sit on per-job networks where a proxy-network container name does not resolve. That call therefore leaves the box and comes back in through caddy. If caddy is not yet listening, the connect is refused and the runner exits 1 rather than retrying in-process.

Because the runner declares no networks, modules/containers/default.nix gives it no after/requires at all. Nothing has ever ordered it behind caddy.

This was known. Commit 9326a8c, 2026-09-13, closes with:

Separately, and NOT fixed here: the runner declares no networks, so modules/containers/default.nix gives it no after/requires at all. Any future deploy that does restart docker.service reproduces this same failure. Worth an ordering dependency, as its own change.

That change was never written. This is the “future deploy”.

5. Was the cause “combining a system update with container updates”?

No — and the timeline is what rules it out. It is worth writing down because it was the first hypothesis, and it is a reasonable one.

The four updates were deployed in four separate deploys, not one. At the moment of failure, deploy 3 contained flake.lock and the inert tofu lock file and nothing else. The three container bumps were already live and untouched; the mailserver major was already live and untouched. Removing them from the branch entirely would have changed nothing about this failure.

What actually mattered is a property of the system update on its own: a new nixpkgs changes the store path of every systemd unit, so switch-to-configuration restarts every container simultaneously. That is the condition the runner cannot survive, and it needs no container bump to arrive. The same thing happened three days earlier for an unrelated reason — commit 523cbe1, enabling the Actions cache, pinned the docker daemon’s bip, which restarted docker.service, which stopped every container — and produced the identical error string.

But the instinct behind the hypothesis is correct, and it is the real lesson. The error was treating a system update as just another row on the same checklist. The dashboard listed lock file maintenance directly beneath update amir20/dozzle docker tag to v11.1.0, as though they were the same kind of change. They are not, and the difference is blast radius:

restartsrebootrollback
a container tag bump1 unitnore-deploy, unless it migrated its own data
lockFileMaintenanceevery unityes, if the kernel movesre-deploy, cleanly

A window sized and sequenced for the first kind is not a window for the second. So: not “don’t combine them in a PR” — they combine in a PR fine, and did. It is “don’t deploy them in the same window, and never deploy the system one last.”

6. What went well

  • kuma-check caught it. The out-of-band timer failed at 22:20:08 and 22:25:01 and was green either side. The self-hosted-status-page paradox described in Observability was closed exactly as designed: kuma could not report its own outage, and the thing that noticed was the probe that lives outside it.
  • deploy-rs did its job on deploys 1, 2 and 4 — three clean activations with magic rollback confirmed, ~40s each.
  • Backups were taken cold and verified before every migrating change. Two local tars and two tagged restic snapshots, all taken with the relevant containers stopped. None were needed. That is the correct outcome.
  • The pre-flight on the mailserver major caught a real break before any downtime: opendkim’s gid moves 104 → 102 in v16, which would have made the box receive mail and silently refuse to send it.
  • The documented recovery was correct. systemctl restart docker was the right first move and it worked. See Failure modes.

7. What went badly

  • A known landmine was left armed for three days. 9326a8c diagnosed the exact failure, wrote down that it would recur, and deferred the fix. The deferral was reasonable in isolation and wrong in aggregate: the cost of the fix was a 59-line module addition, and the cost of not doing it was a six-minute full outage.
  • The system update was attempted in the last third of the window, after it had already been identified as needing its own. Thirteen minutes was not enough for a three-week nixpkgs move plus a reboot plus verification, and the window should have been re-scoped rather than squeezed.
  • deploy-rs’s abort path leaves the box worse than either config. This is the most surprising finding here and deserves its own line: when activation fails, the de-activation stops the containers and the rollback does not restart them. The box does not return to the old generation’s running state — it returns to the old generation’s configuration, with nothing running. A 3-second transient in one non-critical unit therefore cost fifteen healthy services.
  • switch-to-configuration switch failed with Could not acquire lock during recovery, because deploy-rs still held it. Recovery had to fall back to starting units by hand. Worth knowing before the next incident.

8. What changed

  • b380c8a fix(forgejo-runner): a forgejo-runner-ready.service oneshot, in the same shape as forgejo-runner-token, that probes /api/v1/version and waits before the runner starts. Ordering alone would not have been enough — a docker-* unit counts as started when the container is created, not when the service inside it answers. It exits 0 on timeout deliberately: its job is to close the race, not to become a new way for a deploy to fail.

    It earned its place on the first cold boot after the reboot:

    curl: (7) Failed to connect to git.hu-tao.dev:443 after 83 ms: Could not connect to server
    forgejo answered after 2 attempt(s)
    

    The first probe failed. Without the gate, that is the runner exiting 1.

  • The handbook’s Failure modes gains a row for the deploy-rs abort behaviour. (Done; it was ARCHITECTURE.md §12 when this was written.)

9. Action items

#ActionWhyStatus
1Audit every container unit for an unexpressed runtime dependency on another container. The runner was found the hard way; it is unlikely to be the only one.Same class of bug, same trigger.Open. The runner was fixed in b380c8a; no other unit has been audited.
2Deploy lockFileMaintenance alone, first in a window, never last.§5.Standing practice.
3Decide whether a failed activation should de-activate at all. magicRollback protects against losing SSH; it is not obviously the right tool for one crashlooping non-critical unit, and its abort is more destructive than the failure it responds to.§7.Open. magicRollback and autoRollback are unchanged in flake.nix.
4Confirm the healthchecks.io grace period is short enough that two missed pings actually page. The probe failed correctly; whether that produced an alert is configured outside this repo.§6 is only half-verified.Unverified; configured outside this repo.
5Create postmaster@hu-tao.dev.RFC 5321 §4.5.1.Done in 0944f97, with abuse@ alongside.