Introduction
This is the whole of hu-tao.dev: a mail server, a Forgejo instance, a music
library, a search engine, two Minecraft worlds, a Discord bot, and the CI that
builds them — described by one flake, on two machines.
The guiding rule, stated once so the rest follows from it: there is no state on the server that this repo does not describe. One flake builds each machine; OpenTofu creates the machines and publishes their DNS. Anything a human would otherwise have to remember to run is a systemd unit instead.
The two machines
flowchart TB
subgraph vps["hu-tao · CX43 · fsn1"]
direction TB
mail["mail · git · music<br/>search · status · minecraft"]
pages_vol[("pages_data")]
end
subgraph runner["forgejo-runner · CX33 · nbg1"]
direction TB
jobs["job containers<br/>podman, one per job"]
cache[("nix store<br/>Actions cache")]
end
runner -- "HTTPS 443 only<br/>fetch jobs, clone, upload artifacts" --> vps
vps -- "ssh 22, for deploys<br/>initiated here, never there" --> runner
vps -. "pulls published artifacts<br/>on each run, hourly as a safety net" .-> pages_vol
classDef trusted fill:#1f6f43,stroke:#0d3a23,color:#fff
classDef hostile fill:#8c2f2f,stroke:#4d1a1a,color:#fff
class vps trusted
class runner hostile
The green box holds everything. The red box is treated as hostile — it runs code from pull requests, so the design assumes a job will one day escape its container and get root there. What that costs is bounded by the split: a rooted runner holds a nix store, a job cache and its own runner token, and can reach exactly one thing on the VPS, over HTTPS, through an allow-list that permits four URL paths. See CI runner isolation.
Both directions matter and only one is symmetric-looking. The runner reaches
git.hu-tao.dev the way any client on the internet does. The VPS reaches the
runner over ssh because it is the deploy host. There is no private network
between them, deliberately — there was one, it had a single member, and
it was deleted.
Where to start
| If you want to | Read |
|---|---|
| understand how a packet reaches a service | Network and trust boundaries |
| know what runs here and how it is reached | The service stack |
| know what a tailnet peer may reach | The tailnet policy |
| ship a change | Deploying |
| fix something that is broken now | Failure modes and recovery |
| install a machine from nothing | Provisioning a machine |
| know why a thing is the way it is | the module comment — every modules/**.nix opens with why it exists and which mistake it guards against |
That last row is not a deflection. This book is the index to those comments, not a replacement for them: the reasoning lives next to the code it constrains, where it cannot rot independently of it.
Conventions used here
- A “gotcha” has a date. Anything described as learned the hard way names when, so a reader can tell a live constraint from a superstition.
- Numbers are measured, not estimated, and say where they came from.
- Diagrams are mermaid, in the markdown. They render in this book, in the Forgejo web UI, and in a pull request that changes one — so a diagram cannot quietly go stale behind a build step.
Hosts and boot
Three configurations, two machines
flake.nix builds three systems from two builders:
| Configuration | Builder | Root disk | Used by |
|---|---|---|---|
nixosConfigurations.vps | mkVps | /dev/vda (virtio) | the local QEMU test VM (nix run .#default) |
nixosConfigurations.vps-hetzner | mkVps | /dev/sda | what tofu/nixos-anywhere installs, and what deploy .#vps activates |
nixosConfigurations.runner-hetzner | mkRunner | /dev/sda | the CI runner box |
The first two differ in exactly one attribute — the disko device — and share
every module, every container, every secret. vps-hetzner is the real system;
vps exists so the same closure can be booted and tested locally.
runner-hetzner is a second nixosSystem in the same flake rather than a
second flake, because it shares boot, hardware, nix, security and the disk
layout verbatim. What it deliberately does not share is listed in
runner/configuration.nix.
Note the argument that is absent. mkVps passes
sops-nix.nixosModules.sops; mkRunner does not. The runner holds no age key
and can decrypt nothing in secrets.yaml, so the module would only add a unit
that fails at boot. This is enforced, not just intended — the flake carries a
check named runner-has-no-secrets, and it is a nix build rather than an
evaluation, because a check that is only evaluated never runs its builder and
its failure branch is inert.
Inputs are minimal and all follows nixpkgs (nixos-26.05): disko (disk
layout), sops-nix (secrets), deploy-rs (the deploy circuit breaker).
Boot and disk
The disk is GPT with three partitions (disk-config.nix): a 1 M EF02
BIOS-boot partition, a 1 G EF00 ESP mounted at /boot, and ext4 root filling
the rest. Both machines use the same layout; the runner’s 80 GB is picked up
without a line changing, because the root partition is size = "100%".
Two hardware facts drive this and are not preferences:
- GRUB, configured for BIOS and UEFI (
boot.nix). Hetzner Cloud boots these VMs in legacy BIOS mode — the running machine has no/sys/firmware/efi. systemd-boot is EFI-only and would install cleanly, then leave an unbootable box on first reboot. GRUB embeds its core image in theEF02partition (BIOS chain-loads it) and writes/EFI/BOOT/BOOTX64.EFI(the fallback UEFI firmware looks for with no NVRAM entry), so the same closure boots either way.canTouchEfiVariables = false, because there is no efivarfs in BIOS mode. - virtio kernel modules in the initrd (
hardware.nix). NixOS’s defaultavailableKernelModulestargets bare metal and contains no virtio drivers. Hetzner presents the disk over virtio, so without these the initrd cannot see/dev/sda, cannot mount root, and drops to an emergency shell — while the provider still reports the serverrunning. The nixos-anywhere--vm-testcannot catch this; the test harness injects its own virtio modules.
The ESP is 1 G, not 512 M, because it is /boot and holds a kernel + initrd
per generation. It cannot be grown without a reinstall, and filling it breaks
the next deploy rather than the current boot.
The machines themselves
hu-tao 163906050 cx43 fsn1 167.233.24.58 mail, git, everything
2a01:4f8:c015:b138::/64
forgejo-runner 166488672 cx33 nbg1 46.225.61.172 CI only
Both carry delete_protection and rebuild_protection, with user_data in
lifecycle.ignore_changes. They are rebuilt in place, never recreated — CX
instance types are limited availability, and a destroy/create cycle can
find nothing to create.
The runners are a fleet rather than a box: for_each over runner_names in
tofu/, with the address and identity maps keyed by server name rather
than by index. An address on its own says nothing about which box it belongs
to, and a bare list invites being correlated by index with some other list —
which stays silently wrong when an entry is removed from the middle.
The VPS’s public IPv4 is its own resource with delete protection, deliberately separate from the server, so an IP handover carries mail reputation to a new box with no DNS change. See Migration off Ubuntu.
Network and trust boundaries
Traffic reaches the VPS through three doors, and each service sits behind exactly one of them.
flowchart TB
net(("Internet"))
tailnet(("Tailnet"))
net --> edge["<b>Door 1</b> · Hetzner edge firewall<br/><code>tofu/modules/hetzner-firewall</code>"]
edge --> nft["<b>Door 2</b> · host nftables<br/><code>modules/firewall.nix</code>"]
tailnet --> pol["<b>Door 3</b> · tailnet policy file<br/><code>tofu/tailscale-policy.hujson</code>"]
pol --> ts["tailscale0<br/>accepted wholesale on input"]
nft --> inp["input hook"]
nft --> fwd["prerouting DNAT → forward"]
ts --> inp
inp --> hostsvc["host namespace<br/>sshd :2222 · grafana :3000<br/>tempo :4317/4318 · pgbouncer :6432<br/>syncthing GUI :8384<br/>tarpit :222/2022/22222"]
fwd --> dock["docker networks<br/><code>proxy</code> · <code>botnet</code>"]
dock --> caddy["caddy"]
caddy --> pub["public vhosts :443"]
caddy --> priv["tailnet vhosts :8443"]
classDef door fill:#2d4a7c,stroke:#16233c,color:#fff
class edge,nft,pol door
The two firewalls
The edge firewall (tofu/) and the host firewall (modules/firewall.nix) are
kept deliberately similar: each is what survives a misconfiguration of the
other. A port opened at one but not the other is still closed. When something
is unreachable, check both.
input vs forward — the single most load-bearing fact
A published container port is DNAT’d in prerouting and then forwarded — it
never touches the input hook. Docker writes its own accepts into the
ip filter table; in nftables every table’s chain runs, and an accept in
docker’s table cannot rescue a packet that table inet nixos-fw drops. So:
- A service’s published port belongs in the forward allow-list, not input. Getting it backwards yields a port the internet can reach that the firewall never authorised.
- Host-namespace services (sshd on 2222, grafana, tempo, pgbouncer, syncthing, the SSH tarpit, prometheus) are on the input hook.
networking.nftables.flushRulesetmust stay false. The default flushes the entire ruleset — including the tables docker owns — on every reload, and docker only rebuilds them whendockerdstarts. The symptom is latent: running containers keep working, the next container start fails withiptables: No chain/target/match by that name, and recovery issystemctl restart docker.
This boundary is also why fail2ban jails for containerised services set
chain_hook = forward, which carries a sharp edge of its own — see
Failure modes.
Docker networks
| Network | Subnet | Purpose |
|---|---|---|
proxy | docker’s pool | caddy resolves its upstreams here by container name over docker’s embedded DNS — no IP addresses in the Caddyfile |
botnet | 172.30.0.0/24 (pinned) | the discord bot + its redis, isolated from the proxy. Pinned because infra.botGateway (172.30.0.1) is a literal in the bot’s DATABASE_URL, its OTLP endpoint, and the firewall’s input rule |
modules/containers/default.nix creates each network as a oneshot systemd unit
that every container on it requires, so a container can never start onto a
network that does not exist yet.
infra.dockerBridgeGateway (172.17.0.1) is a third address that matters: it is
how a container reaches a service in the host’s network namespace, which is
what caddy does for grafana:3000 and syncthing’s GUI:8384. It is observed,
not enforced — deliberately not pinned with the daemon’s bip setting, even
though pinning is what botSubnet/botGateway do for a network this repo
creates itself. bip is daemon config, so setting it restarts docker.service,
which stops every container. A wrong value here costs one failed container
start and a rollback; pinning it costs a full container restart on every deploy
that touches the line.
The input rule that makes that route work is
iifname "br-*" ip saddr != 172.30.0.0/24 tcp dport { 3000, 8384, 8443 }, and
the exclusion is the point. The rule above it grants the bot exactly two ports
(4317, 6432) from exactly botSubnet; without the !=, this one immediately
handed the same bridge three more, so a narrow grant was followed by a broad one
and only the broad one meant anything. The bot is the right container to
subtract first — it is the only one here whose input is arbitrary text from
strangers that it then ships to a third-party model.
What remains matched is deliberate and worth knowing: searxng, kuma,
forgejo, navidrome, dozzle, cloudflared, mailserver, webmail and pages-hook
share the proxy bridge with caddy, so they are still admitted to those three
ports. Both services behind them are credential-protected (grafana has a real
admin login with sign-up off; syncthing’s GUI password is declared by
modules/syncthing.nix through guiPasswordFile and re-applied on every
deploy), so this is defence in depth, not a hole being closed.
That syncthing half used to be an assertion about the live box rather than
something the deploy enforced. The config directory came across from the old
host with a password already set by hand, and nothing in the repo would have
noticed its absence: on fresh state syncthing comes up with a generated API key
and no password at all, bound to 0.0.0.0 and reachable from every bridge
in the list above. That mattered more than a login form normally does, because
syncthing’s REST API is not a viewer — /rest/config/folders takes a
versioning block of type external with a params.command that syncthing
execs, and a folder rooted at ~/.ssh writes authorized_keys as hutao, who
has passwordless sudo. The credential is now declarative, so it fails closed.
Subtracting the rest needs a positive source match, which needs a
pinned subnet on proxy, which means deleting a network that already exists and
detaching every container on it — a maintenance window, not an edit.
One silent consequence of adding ip saddr: the rule is now IPv4-only,
where the bare version matched both families. That is free today because
docker’s bridges here carry no IPv6, but turning on docker IPv6 would need an
ip6 saddr != sibling or those three ports go dark over v6 with every other
check still passing.
Tailscale
The tailnet interface is accepted wholesale on the input hook, which is how
the tailnet-only services (grafana, tempo, dozzle, pgbouncer, syncthing GUI)
are kept private — by the absence of an internet rule, not by their bind
address. Several bind 0.0.0.0 and rely entirely on this.
That is also why door 3 is a policy file and not an nftables rule. nftables
cannot tell tailnet peers apart: every packet off tailscale0 looks the same
to it, so “which peers may reach what” is a question this layer is structurally
unable to answer. tofu/tailscale-policy.hujson is the layer that can, and it
narrows the VPS to ten ports for the owner account — 6432 and 4317/4318 are
no longer tailnet-reachable at all. See
The tailnet policy.
Tailscale SSH is off (tailscale set --ssh=false). With it on, tailscaled
owned port 22 on the tailnet address, and once split DNS sent git. there on
tailnet devices, git over ssh landed on Tailscale SSH instead of forgejo.
Administration, deploys and the runner’s jump hop all use sshd on 2222 with an
ordinary key.
Three of those services also answer by name — dozzle., grafana. and
syncthing. — and that is the same mechanism wearing a hat. Caddy runs a
second listener on infra.tailnetHttpsPort (8443) carrying those three
vhosts, plus tailnet copies of git. and status. with their admin paths
open (see split DNS); like
every other private port it is published on 0.0.0.0 and kept private by being
in neither allow-list.
What makes the URL portless is a nat prerouting chain:
type nat hook prerouting priority -110; policy accept;
iifname tailscale0 tcp dport 443 redirect to :8443
iifname tailscale0 tcp dport 80 redirect to :8880
iifname tailscale0 udp dport 443 redirect to :8443
The udp line is HTTP/3. The tailnet listener speaks it too (8443/udp is
published), and its sites advertise h3=":443" rather than caddy’s default
h3=":8443": udp 8443 sent directly over the tailnet did not get through when
measured, while udp 443 through this redirect did. A browser that learned the
public listener’s h3=":443" takes the same path once split DNS points git.
or status. at the tailnet address. Without the redirect that traffic hung —
Forgejo’s webpack chunk loads failed on it — and with docker’s rule it would
reach the public listener, whose copies of those names close the admin paths.
Known quirk, unexplained: udp 8443 sent directly over the tailnet does not
reach caddy, while tcp 8443 direct and udp 443 through the redirect both do.
Nothing uses it — the tailnet sites advertise :443 — so it is recorded here
rather than chased. docker port caddy and nft list ruleset | grep 8443 on
the box are where to start if it ever matters.
Two details in those lines do all the work, and they are independent:
iifname tailscale0is what keeps this off public traffic. A request arriving on the public interface never matches, so it falls through to docker’s own prerouting and reaches caddy’s ordinary:443listener exactly as it always did. This chain cannot affect a public vhost.- priority
-110is ten ahead of docker’sdstnatat-100, and that ordering is the entire control. Both chains run on the same hook; the lower number runs first.
flowchart TB
pkt["tcp dport 443"] --> iif{"which interface?"}
iif -- "public" --> dn["docker dstnat -100<br/>ours never matched"]
dn --> pub443["caddy :443<br/>public vhosts"] --> ok1(["200"])
iif -- "tailscale0" --> prio{"our chain's<br/>priority?"}
prio -- "-110, ahead ✓" --> rw["redirect 443 → 8443"]
rw --> priv["caddy :8443<br/>tailnet vhosts"] --> ok2(["200"])
prio -- "after -100 ✗" --> dn2["docker dstnat first"]
dn2 --> pub443b["caddy :443<br/>public vhosts"] --> nf(["404<br/>no such site"])
classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
classDef good fill:#1f6f43,stroke:#0d3a23,color:#fff
class nf,dn2,pub443b bad
class ok1,ok2,rw good
The right-hand branch is the misconfiguration, not a second real path — there is no case in which a tailnet request legitimately lands on the public listener. It is drawn because getting the priority wrong fails silently: the packet is still delivered, caddy still answers, and the only symptom is a 404 on a name that resolves, from a box that is up, over a link that works.
The names are plain A records to the box’s 100.x address
(tofu/modules/cloudflare-dns). A CNAME to the node’s MagicDNS name would
avoid that literal — the trap modules/containers/tempo.nix documents — but
Tailscale does not publish <node>.<tailnet>.ts.net in public DNS (verified
2026-09-17: empty answers from 1.1.1.1, 9.9.9.9 and 8.8.8.8), so it resolves
only on a device whose MagicDNS is active and fails silently on one where it is
not. The literal is the lesser failure: it goes stale only when the machine is
replaced, and it goes stale loudly. The certificate covers the names as
ordinary SANs, because DNS-01 never asks whether a name resolves publicly.
Why there is no private network
The two machines had a Hetzner private network. It was deleted on 2026-09-19, and the reasoning is worth keeping because the shape recurs:
- It had one member. A private network with a single host is a subnet, not a topology.
- hcloud firewalls do not filter private traffic. Door 1 simply does not exist on that interface.
- The VPS input chain accepts nine ports with no
iifnamequalifier, so every one of them was reachable from the private interface as readily as from the internet.
Together that is a standing bypass of two of the three doors, waiting for a second member to make it exploitable. The runner reaches the VPS over ordinary public HTTPS instead, where all three doors apply and Caddy can read the request. See CI runner isolation.
The tailnet policy
The tailnet was, until 2026-09-20, the one piece of infrastructure this repo
did not describe. modules/firewall.nix accepts iifname tailscale0
wholesale — that is how every tailnet-only service is kept private, by the
absence of an internet rule rather than by a bind address — so what a tailnet
peer could reach was decided entirely in a web console, by a document nothing
here could review, diff or roll back.
The consequence was a flat trust zone. Anything on the tailnet reached sshd on
2222, pgbouncer on 6432, tempo’s two OTLP receivers, grafana, dozzle and
syncthing’s GUI. One compromised phone was the whole estate. nftables cannot
tell tailnet peers apart; the policy file is the only layer that can, and it
is now tofu/tailscale-policy.hujson.
One document, no partial ownership
The obvious wish is for tofu to own some rules and leave the rest to the
console. The provider forecloses it: tailscale_acl “controls a tailnet’s
entire policy file and not just the ACLs section within it”, and it “will
completely overwrite existing policy file contents”. There is no section-level
ownership, and no ignore_changes that could give it — the whole document is
one string attribute.
So anything omitted from that file is deleted on apply, and the plan does not
say so. The first real plan here carried a nodeAttrs block granting four
devices Mullvad exit-node access plus tailnet Funnel, and a second tagOwner
with its ssh rule — none of it in the first draft, all of it silently on its
way out. It is carried over verbatim now, with a comment saying why each line
cannot be tidied. Before any apply that replaces this resource, read the live
policy in the admin console and diff it by eye.
HuJSON comments survive the apply and show up in the console, so the file is also the documentation the next person reads there.
Tags are the only way to narrow anything
This is the part that decides the shape of the whole file: Tailscale has no deny rules. Rules are purely additive. You cannot narrow a destination by adding a rule — you can only stop a wider rule from covering it.
The wider rule is the stock one, device-to-device over autogroup:self. And
per Tailscale’s own documentation, “autogroup:self only applies to user-owned
devices. It does not apply to tagged devices.” So the moment the VPS carries a
tag it drops out of that rule, and what is left is the explicit port list.
flowchart TB
peer(("a tailnet peer"))
peer --> pol{"policy file"}
pol -- "autogroup:self:*<br/>every port" --> own["your own devices<br/>laptop · desktop · phone"]
pol -- "tag:vps<br/>ten ports" --> vps["the VPS"]
pol -- "tag:friends-ssh:22" --> friends["a friend's machine"]
vps --> nft["host nftables<br/>iifname tailscale0 accept"]
nft --> svc["sshd:2222 · grafana:3000<br/>dozzle:8080 · caddy:8443"]
gone["pgbouncer :6432<br/>tempo :4317 / :4318"]
pol -. "no rule names these" .-> gone
classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
class gone bad
Tagging is deliberately not in the file. Applying a tag to a device is a
console action, or tailscale up --advertise-tags. So moving a machine between
trust classes needs no apply, no commit and no credential — which is the part
that genuinely wanted to be ad-hoc — while the rules themselves change rarely
and go through review.
What is live
acls: 1git2clone@github -> autogroup:self:*
1git2clone@github -> tag:vps:22,53,80,443,2222,3000,8080,8384,8443,8880
1git2clone@github -> tag:friends-ssh:22
ssh: check -> autogroup:self as nonroot, root
accept -> tag:friends-ssh as nonroot, root
Every src is the owner account rather than autogroup:member, and that is a
smaller change than it looks. autogroup:self as a destination means
“devices belonging to whoever opened this connection”, evaluated per
connection — so member A could never reach member B’s laptop through it. What
naming the account removes is the one remaining case, a second member reaching
their own devices. On a tailnet with one real user that is a no-op today and
a closed door the day it stops being one.
The port list on tag:vps is everything a tailnet client legitimately reaches:
| Port | What |
|---|---|
| 2222 | the host’s own sshd — deploy-critical, deploy .#vps connects here. Drop it and deploys stop |
| 22 | forgejo’s ssh, for git over the tailnet |
| 53 | the split-DNS resolver: answers git. and status. with this box’s tailnet address, see below |
| 80 / 443 | redirected to 8880 / 8443 by the firewall’s prerouting. The ACL sees the port the client sent, before the rewrite, so these are the ones that must be listed |
| 3000 | grafana, host networking |
| 8080 | dozzle direct — the path that still works when caddy is the broken part |
| 8384 | syncthing’s GUI |
| 8443 / 8880 | caddy’s tailnet listeners, reachable directly as well as through the redirect |
Deliberately absent: 6432 and 4317/4318. pgbouncer and tempo’s OTLP
receivers bind 0.0.0.0 and are reached by the discord bot over the docker
bridge, not from the tailnet — the input chain has a separate rule for that,
scoped to botSubnet. Nothing on the tailnet has ever needed them. A tailnet
peer reaching postgres is the shape of the incident this file exists to make
impossible.
tag:friends-ssh is a destination and never a source
The tag predates this file. A friend’s machine carries it for exactly one
reason: so it can be SSHed into without the rest of the tailnet’s members
reaching it. Tagging it took it out of its owner’s autogroup:self — the same
lever tag:vps uses.
One tag for the class, not one per machine. A second friend’s laptop joins this tag rather than arriving with a tagOwner, three rules and two tests of its own that a reviewer would have to read in full to discover they grant exactly what this one already grants.
Assigning it is a console action, and the order matters the same way it does
for tag:vps: a device cannot be given tag:friends-ssh until an applied
policy defines it, so tofu apply comes first and
Machines → the device → Edit ACL tags second. A device whose tag no rule
here names has no grant at all — autogroup:self does not cover it either,
because it is tagged.
Nothing on those machines has any business opening a connection into this
tailnet, and since there are no deny rules, “cannot initiate” is not a rule you
write. It is what you get by never naming the tag as a src. The stock policy
named it implicitly, because src: ["*"] covers tagged devices; deleting that
wildcard is the whole fix.
Which produced the least obvious line in the file. A Tailscale SSH rule is an
extra check layered on top of ordinary ACL matching, not a substitute for it —
the connection must still be permitted to reach port 22 on the destination. The
old dst: ["*"] granted that for free; with the wildcard gone,
tag:friends-ssh:22 has to be granted explicitly in acls or ssh to those
boxes simply stops working,
and the ssh section looks innocent while it does. A test asserts it.
What check actually checks
The name invites a wrong guess. check does not verify that the connecting
user owns the destination — that decision was already made by src/dst
matching. What it adds is freshness: the connecting user re-authenticates
to their Tailscale identity in a browser, and the answer is cached for about
12 hours.
So ownership comes from autogroup:self and from naming the account in src.
check is what makes a stolen but still-enrolled laptop not be enough on its
own. It covers your own devices only: Tailscale SSH is off on the VPS, which
is administered over sshd on 2222 — see Deploying.
The tests are the part worth having
The policy above is a claim; the tests block is that claim being checked, by
Tailscale, against the real tailnet:
accept vps:2222 · vps:443 · vps:53 · vps:8080 · tag:friends-ssh:22
deny vps:6432 · vps:4317 · tag:friends-ssh:80
A rule that stops being true fails the apply. vps:2222 is there because it is
the single most load-bearing line in the file, and tag:friends-ssh:22 because
it catches exactly the “tidy-up” described above.
They run at apply time, not plan time — measured, not assumed. An earlier
draft of that comment said plan. A real tofu plan then succeeded against a
policy carrying both an empty vps host and impossible assertions, which means
the provider diffs the string locally and ships it to the API only on apply. So
a broken policy reaches tofu apply before anything objects.
The provider then reports test(s) failed (400) and swallows the detail.
To see it, POST the rendered file to /api/v2/tailnet/-/acl/validate with an
OAuth token — that endpoint validates what you hand it and installs nothing:
[acl test error]: address "vps:6432" (protocol "tcp"): want: Drop, got: Accept
[acl test error]: address "vps:4317" (protocol "tcp"): want: Drop, got: Accept
That message is not a bug to work around. It is the file saying the narrowing is not live yet, because the VPS is not tagged.
The bootstrap deadlock
Which it was, once, and unavoidably:
- The policy cannot be applied while the VPS is untagged, because
autogroup:self:*still covers it on every port and the deny tests say otherwise. - The VPS cannot be tagged while the policy is unapplied, because Tailscale
refuses to assign a tag that no
tagOwnersentry defines — andtag:vpsis defined only in the policy waiting to be applied.
Neither the console nor a tailscale_device_tags resource gets around that; it
is the API’s rule, not a tooling limit. var.vps_is_tagged breaks the cycle by
dropping exactly those two assertions for exactly one apply:
tofu apply -var vps_is_tagged=false # publishes tagOwners; changes no access
# Machines -> vps -> Edit ACL tags -> tag:vps (now offered)
tofu apply # the apply that actually narrows anything
It defaults to true, so the un-narrowed policy is never what you get by
accident — reaching for the weaker one has to be deliberate and visible in the
command. With a saved plan the variable goes on the plan, not the apply:
tofu plan -var vps_is_tagged=false -out main.tfplan.
This is a three-step sequence once in the life of a tailnet, not a routine.
The VPS joins tagged, not tagged afterwards
Tagging by hand left a gap that no plan would show: a VPS rebuilt from scratch
rejoins as an ordinary user-owned device, which autogroup:self:* covers on
every port. That is a silent return to the flat tailnet, visible only at
the next tofu plan — and deploy .#vps does not run tofu.
So the tag is now half of the credential. modules/services.nix passes
--advertise-tags=tag:vps, which is not optional: Tailscale refuses to register
a device with an OAuth client secret untagged.
sops.templates."tailscale-authkey".content =
"${config.sops.placeholder.tailscale_oauth_client_secret}?ephemeral=false&preauthorized=true";
An OAuth client secret, not a tskey-auth- key, because the auth keys it
replaces capped out at 90 days — the box was one forgotten rotation away from
being unable to rejoin its own tailnet after a rebuild.
The two query parameters are in the clear on purpose, and one of them is
load-bearing in a way that is invisible when wrong. An OAuth-minted key is
ephemeral=true by default, and an ephemeral node is removed from the tailnet
when it goes offline. Left at the default, this VPS would delete itself on
every reboot and come back as a new node with a new tailnet address — so the
dozzle/grafana/syncthing A records and the vps host in this policy would
both point at nothing. Buried inside ciphertext, that is a one-character mistake
nobody can review. preauthorized=true only matters if device approval is ever
turned on, and is the difference between an unattended rebuild and one that
waits for a human to click approve.
Changing the flags is safe on a running node: tailscaled-autoconnect only runs
tailscale up when the backend state is NeedsLogin, NeedsMachineAuth or
Stopped. On a node that is already Running it exits without touching
anything, so this takes effect on a fresh join and changed nothing the day it
landed.
Which means the live node’s tag does not come from that flag yet, and the
check for it is not the obvious one. tailscale debug prefs reads
"AdvertiseTags": null on this box today, because the tag was applied in the
console during the bootstrap above and the node has never re-run tailscale up
since. The tag is nonetheless set, server-side, which is what the policy
evaluates against. Ask the control plane, not the local prefs:
$ tailscale whois 100.109.115.12
Machine:
Name: vps.dikdik-cloud.ts.net
Addresses: [100.109.115.12/32 fd7a:115c:a1e0::493b:730d/128]
Tags: tag:vps
The two agree only after a rejoin. Until then --advertise-tags is insurance
against the rebuild, not a description of the present state.
The OAuth client wants one scope — Keys → Auth Keys → write — and tag:vps
selected, since a client can only mint keys for the tags chosen at creation.
The client tofu uses for the policy itself is a different one, scoped to
Policy File (write).
What this still does not cover
nodeAttrsis carried, not understood by this repo. The four addresses grantedmullvadare devices with the integration enabled; dropping one takes that machine’s Mullvad exit nodes away with nothing in the apply output saying so. They are raw addresses because that is how the console wrote them.- The tailnet has one real user. Every
srcnames that account, so a second member gets nothing until a rule says otherwise — which is the right default and also means adding a person is a policy change, not an invite. - nftables still accepts
tailscale0wholesale. That has not changed and should not: it is the layer that cannot tell peers apart. The policy file is now the layer in front of it, and the two are the usual arrangement here — each is what survives a misconfiguration of the other.
Split DNS for the half-public names
git.<domain> and status.<domain> are served twice: publicly, and on caddy’s
tailnet listener with their admin paths open (see
Access control). The
tailnet-only names need nothing like this — their public A record already
points at the tailnet address, which only the tailnet can reach. These two must
keep their public address for everyone else, so one public record cannot serve
both audiences.
Tailscale split DNS closes it: tofu’s tailscale_dns_split_nameservers sends
lookups for exactly those names, from tailnet devices only, to a CoreDNS on the
vps bound to its tailnet address. It answers them with that address and refuses
everything else. Nothing is configured per device, phones included.
- The name list is computed in
modules/containers/caddy.nix— every host that has both a public and a tailnet copy — and tofu’svar.split_dns_subdomainsmust agree with it. - It needs port 53 in the grant above, and the DNS (write) scope on the OAuth client alongside Policy File.
- HTTP/3 has to follow the same route. A browser that saw the public
Alt-Svc: h3=":443"keeps it for 30 days and then sends QUIC to udp 443 at the tailnet address. The firewall redirects tailscale0’s udp 443 to the tailnet listener, which speaks HTTP/3 as well — see Network. Before that redirect existed, those requests hung and Forgejo’s webpack chunk loads failed. - If the resolver is down, your tailnet devices most likely cannot resolve
those two names at all. Split DNS sends them only there and Tailscale
documents no fallback. Everyone else is unaffected, and both names live on
this same box anyway, so the case that matters is CoreDNS alone failing,
which systemd restarts. Order matters for the same reason: deploy the
resolver before
tofu applypoints the tailnet at it.
CI runner isolation
The runner is the one machine here that runs code this repo did not write. Every other design decision in this book protects a service from the internet; this one protects the estate from its own CI.
What it replaced
The runner used to be a container on the VPS with /var/run/docker.sock
bind-mounted in. That socket is root on the machine serving mail, git, every
sops secret and two Minecraft servers — so any workflow, including one opened
by a dependency bot, was one docker run -v /:/host away from the entire
estate.
The isolation boundary did not get stronger. It moved — from a container to a VM — and what sits inside it got much smaller.
Seven layers
flowchart TB
subgraph R["forgejo-runner · hostile"]
direction LR
job["job container<br/>podman, per job"]
rd["runner daemon"]
end
job --> L4
rd --> L4["④ runner nftables<br/>out: policy-drop, VPS on 443<br/>in: ssh from the VPS only"]
L4 --> L1["① runner cloud firewall<br/>out: 53 / 80 / 443 · udp 123"]
L1 --> L2["② VPS cloud firewall<br/>in: tcp/443 from anywhere"]
L2 --> L3["③ VPS nftables"]
L3 --> L5["⑤ caddy L7 allow-list<br/>4 paths, everything else 403"]
L5 --> cad
subgraph V["hu-tao"]
direction LR
cad["caddy"] --> fj["forgejo:4242"]
end
classDef ctl fill:#2d4a7c,stroke:#16233c,color:#fff
class L1,L2,L3,L4,L5 ctl
| # | Control | Where |
|---|---|---|
| 1 | Runner cloud firewall: in = tcp/22 from the VPS /32 only; out = 53/80/443 + udp/123 (NTP) | tofu/runner-firewall.tf |
| 2 | VPS cloud firewall: tcp/22 outbound to the runner /32 | tofu/modules/hetzner-firewall |
| 3 | VPS nftables: output chain is policy-drop, one rule per runner address | modules/firewall.nix |
| 4 | Runner nftables: output policy-drop, the VPS reachable on 443 and nothing else; inbound ssh accepted from the VPS alone | modules/runner/firewall.nix |
| 5 | Caddy: runner addresses restricted to four Actions paths, everything else 403 | modules/containers/caddy.nix |
| 6 | Job: no engine socket by default; its container is created per job and destroyed after | modules/runner/default.nix |
| 7 | pages-pull: a fetched artifact reduced to files and directories, modes normalised, before caddy serves it | modules/pages-pull.nix |
Every connection between the two hosts is either initiated by the VPS, or
is HTTPS from the runner to git.hu-tao.dev like any other client on the
internet. There is no private link, deliberately —
see why.
Layer 7 is a different kind of control from the six above it, which is why it was missing for a while. Every one of those reads an address, a port or a path; none of them can read what is inside a request that the allow-list legitimately permits. The published artifact is exactly that — content the runner authors, travelling through a route layer 5 has to admit, landing in a tree caddy serves. What it strips, and why.
Who may ssh in
Layers 1 and 4 both narrow inbound ssh to the VPS’s /32, and that duplication
is the point — it is the same “each survives the other’s misconfiguration”
arrangement as the two firewalls on the VPS. The host rule is:
ip saddr <vps4> tcp dport 22 ct state new counter name ssh_from_vps accept
tcp dport 22 ct state new counter name ssh_blocked log prefix "DROP_ssh: " drop
It used to be a bare tcp dport 22 ct state new accept, on the reasoning that
only the cloud firewall should key on the VPS’s address, since a host rule would
have to survive that address changing. The rest of the file had already taken
that bet: the output and forward chains key their one-way drops on the same
address, there is an assertion that it is non-empty, and the VM test exists
precisely to override it.
What settles it is that the two failure modes are not symmetric. A stale address in the egress rules fails open — the drop matches nothing, the runner reaches the new address on every port, and the one-way design is gone with no error anywhere. Stale here fails closed: nobody can ssh in, Hetzner’s web console still works, and the fix is one rebuild. Closed is the direction to be wrong in.
IPv4 only, matching the cloud rule, which lists a v4 /32 and no v6 source at
all — so v6 ssh was never reachable through the layer above. The second line is
redundant with the chain’s drop policy and exists to be named: the catch-all
carries an anonymous counter, so without it a refused ssh is indistinguishable
from any other dropped packet.
Verified live after the deploy, from the VPS:
$ ssh -J vps root@46.225.61.172 nft list counter inet nixos-fw ssh_from_vps
counter ssh_from_vps {
packets 2 bytes 120
}
The one-way rule, and how it is enforced
The runner’s nftables output chain is policy-drop, and the three lines that
matter are ordered above every broad accept:
ip daddr <vps4> tcp dport 443 ct state new counter name vps_allowed_out accept
ip daddr <vps4> counter name vps_blocked_out log prefix "DROP_vps_out: " drop
ip6 daddr <vps6> counter name vps_blocked_out6 log prefix "DROP_vps_out6: " drop
Order is the whole control. nftables is first-match-wins within a chain,
and the chain below these carries oifname "podman*" accept and
tcp dport { 53, 80, 443 } accept. Move the drop under either one and a
generic tcp dport 443 matches first, so the drop never runs — and nothing
visibly breaks, because the traffic it was supposed to stop is traffic that
also works. The same three lines are repeated in the forward chain, because
job containers route through it rather than through output.
The v6 line drops the VPS’s whole /64 outright: nothing legitimate goes there
over v6, and a runner that can reach the box on any v6 address has defeated the
v4 rules.
The metadata service is the host’s alone
The forward chain carries one drop the output chain deliberately does not:
iifname "podman*" ip daddr 169.254.0.0/16 counter name metadata_blocked_fwd drop
169.254.169.254 answers this server’s own user_data, and tofu/server.tf
puts the box’s Forgejo registration pair there — the per-runner entry from
var.runner_identities.
modules/runner/identity.nix reads it once at boot from the host’s netns,
which is why the output chain allows tcp/80 and this drop leaves that path
alone.
A job container is a different netns, so its packets are forwarded and land
here instead. Without the rule they reach the same endpoint, and one curl in
a workflow returns the uuid and secret — enough to register a second runner
daemon against the instance, call FetchTask, and receive other jobs along
with their tokens. That is precisely the persistence container.docker_host: "-" was chosen to deny, arriving by a route that never touches a container
socket.
It matches the whole 169.254.0.0/16 rather than the single address: nothing a
job does has business in link-local space, and a provider that moves its
endpoint within that block does not get to reopen this quietly.
Not live on the current box. It predates the mechanism and tofu holds
user_data in ignore_changes, so the endpoint returns 204 today and the
identity instead sits in a 0700 root-owned file a job cannot reach. The rule
is there for the next runner tofu creates, which
tofu/variables.tf documents as the ordinary
way to add one.
Two things test this:
tests/runner-firewall.nix— a three-node NixOS VM test with real packets, six subtests. The third node exists only to be not the VPS: without it the ingress half could show that ssh from the VPS is accepted, which was never the half in doubt. It asserts on the counters rather than onnc’s exit status, because no sshd is listening on the test node and a permitted connection is refused exactly like a filtered one is. It needs/dev/kvmand the VPS does not have it (shared-vCPU Hetzner instance, no nested virtualisation), so it is hand-run on a machine with KVM:nix build .#checks.x86_64-linux.runner-firewall -L.runner-firewall-ordering— greps the evaluated ruleset for line order. No KVM, costs seconds, runs in CI. It catches the exact regression the VM test exists for — the VPS rules sinking below a broadpodman*accept — just not with a real packet.
Measured, not assumed: 443 to the VPS returns 200; 22 and 2222 are dropped, with the drop counter moving 0 → 14.
Why layer 5 exists
Layers 1–4 are address-and-port controls. They can say “that box may reach
tcp/443 here” and nothing finer — and the runner must reach 443, because
that is how it fetches jobs. Without something reading the request, a rooted
job gets the whole Forgejo surface: every repo it can see, the web UI, all of
/api/v1.
Caddy narrows that to four paths:
/api/actions/*
/twirp/github.actions.results.api.v1.ArtifactService/*
/*/*/info/refs
/*/*/git-upload-pack
That set was measured from a real run’s access log, not guessed. No
/api/v1, no web UI, and — the one worth saying out loud — no
git-receive-pack: the runner can clone and cannot push. That converts the
open-ended risk “a compromised runner can push to repos it built” into
something a packet filter could never express, because push and fetch share a
port and a TLS session.
/api/actions_pipeline/* was in this list and was deliberately removed. It is
the v3 artifact API, added defensively on a guess that uploads rode it; a full
run’s capture never touched it. An allow-list entry nothing uses is surface,
and this is the layer whose whole job is to have less of it.
remote_ip is the TCP peer and never a header: trusted_proxies is unset and
every DNS record is proxied = false, so there is nothing in front of Caddy to
launder an address. A rooted runner can forge any token it holds; it cannot
forge its source address.
Verified from the runner: 403 on /, /api/v1/version, /explore/repos and
/user/login; allowed paths proxy through; ordinary clients still 200
everywhere; a full workflow run produced no 403 at all.
The job’s engine socket
container.docker_host is -: no engine socket is mounted into the job.
Podman’s socket is root on this box, so a job holding it could plant units or
read the runner’s identity pair, and that would outlive the job. Without it a
compromise lasts one job. services: still work, because the runner creates
service containers, not the job (a9bf861, and the comment above
docker_host in modules/runner/default.nix).
Podman rather than docker, and not as a preference: docker’s nftables
integration is what produced the half-working published ports and forward-chain
traps documented at length in modules/firewall.nix. Jobs still get a working
docker command (dockerCompat plus dockerSocket), so a workflow that
shells out to docker needs no edit.
forgejo-runner 13.1.0’s daemon has no --ephemeral/--once flag, so
one-job-per-runner-lifetime is not available upstream. Per-job disposability
comes from the container being created and destroyed per job instead.
The job image, and the second label
Every JavaScript action — actions/checkout included — is executed by a node
binary inside the job container, not by the runner. The nix label is
nixos/nix, which carries nix, bash, gitMinimal, curl and coreutils and no
node, so on that label uses: cannot run at all and every workflow hand-writes
its checkout as a git fetch.
That is a correctness trap and not only an inconvenience. A hand-written fetch
is anonymous unless the author remembers to thread the job token through it —
and on a public repo an anonymous fetch works. The omission is therefore
invisible on five of the six repos on this instance and fatal on the sixth:
cv-template, the one private repo, failed with could not read Username for 'https://git.hu-tao.dev'. actions/checkout defaults its token input to the
injected job token, so on an image with node the private-repo case needs no
thought from the workflow author.
So there is a second label, and it is additive:
| Label | Image | node | Notes |
|---|---|---|---|
nix | nixos/nix:2.35.2 | no | what this repo’s own workflows run on |
nix-node | localhost/forgejo-ci-nix-node | yes | nix and node; built here, not pulled |
ubuntu-latest | node:22-bookworm | yes | no nix, and no sudo — see CI |
node-22 | node:22-bookworm | yes | |
alpine | alpine:3.22 | no |
nix is untouched, so vps, skavex, compress and serenity-discord-bot
keep building on exactly the image they build on today; a repo opts in by
changing runs-on. A broken image here cannot take CI down for four working
repos.
Why it is not in a registry
force_pull is false, so the runner uses a locally present image without
reaching for a registry. That lets the image be an ordinary Nix derivation
(modules/runner/ci-image.nix) loaded into podman at activation — pinned by
flake.lock, rebuilt only when its inputs change, with no push credential, no
pull secret and no package visibility to get wrong. The label’s reference comes
from config.runner.ciImageRef rather than a repeated string literal, so the
label and the image it names cannot drift apart.
buildLayeredImageWithNixDb, not buildLayeredImage. The plain builder copies
the store paths in but leaves /nix/var/nix/db empty, and a nix that cannot
read its own database treats every path in the image as absent — nix develop
then rebuilds a closure that is already sitting on the disk.
The image provides /usr/bin/env and nothing else from the FHS. dockerTools
links contents into /bin and creates no /usr at all, while nixos/nix
ships the usual env — so a workflow that worked on the nix label dies here
on the most common shebang in the ecosystem, #!/usr/bin/env node. Every binary
npm and pnpm install starts that way, which is why pnpm check on skavex
failed at its first step with a message naming the interpreter rather than the
script.
The garbage collector eats it
podman system prune -f --all runs daily, and --all removes every image no
container references. Between jobs nothing references this one, so it is deleted
like any other cold image — and every job on nix-node then dies in about three
seconds on a pull of a localhost/ reference no registry can serve. Observed
2026-09-21 across cv-template, skavex and nixos-dotfiles at once, on
workflows that were green hours earlier and had not changed.
Two things were wrong, and the second is the one worth carrying forward:
- Nothing put the image back.
podman-prunenow carries anExecStartPostthat restarts the load unit.ExecStartPostrather thanOnSuccess=, becauseOnSuccessonly starts a unit — which is the same trap one layer up. - The load unit had
RemainAfterExit. systemd therefore believed it active forever after its first success and would not run it again; and because the unit’s own text does not change when the image does, a deploy could not restart it either. So the deploy that installed the prune hook fixed the next deletion and left the current one in place, and every job kept failing until the unit was restarted by hand. Dropping the flag is what makes both healing paths work, since activation starts wanted-but-inactive units. It only ever bought skipping apodman loadwhose layers were already on disk.
Identity, and why it is not a Nix string
The runner is a declared runner: a uuid+secret pair Forgejo issues for one
record, written into server.connections in config.yaml. No imperative
forgejo-runner register call (deprecated upstream), no .runner state file.
But this box is a snapshot cloned into N runners, so the uuid cannot be a Nix string the way it was for the VPS’s single permanent runner — every clone would claim the same runner record, which is undefined.
sequenceDiagram
participant H as metadata
participant I as identity.nix
participant D as daemon
H->>I: user_data
I->>H: live instance-id
Note over I: must match, or stop
Note over I: prepend server: to config.yaml
I->>D: start
Note over D: read, declare, poll
Identity arrives via Hetzner user-data and is checked against the live
instance-id before the pair is read, so a cloned snapshot is inert
elsewhere. The static half of config.yaml (capacity, cache, engine settings)
is a Nix-checked writeText; only the genuinely per-instance part is
imperative.
Cache
The Actions cache is served by the runner itself on infra.cacheProxyPort
(34567), reached from job containers over the per-job podman bridge and
admitted by one input rule. The port is fixed rather than random because a
firewall rule cannot name a port the daemon chooses at startup.
A new box starts empty, so the first run after provisioning pays full compile.
serenity-discord-bot, same commit, back to back:
| job | cold | warm | |
|---|---|---|---|
test (postgres + redis services) | 13m2s | 4m20s | 3.0× |
build (--all-features) | 6m43s | 1m32s | 4.4× |
build (--features "opentelemetry ai-openrouter util-download") | 6m34s | 1m32s | 4.3× |
build () | 4m48s | 1m19s | 3.6× |
clippy (--all-features) | 5m30s | 1m15s | 4.4× |
clippy (--features "opentelemetry ai-openrouter util-download") | 4m49s | 1m14s | 3.9× |
clippy () | 4m34s | 1m7s | 4.1× |
fmt — no cache, control | 48s | 40s | — |
fmt is the control: it uses no cache and did not move, which rules out “the
new box is just slower” as an explanation for the cold column.
Why the cache works when artifacts did not
Both run the same hostname test. @actions/* checks GITHUB_SERVER_URL
against GITHUB.COM, *.GHE.COM, *.LOCALHOST, else “GHES” — and on a
Forgejo instance that always says GHES. What each package does with the
answer is opposite:
@actions/artifactv2+ throwsGHESNotSupportedErrorbefore opening a socket.actions/upload-artifact@v4therefore fails with zero HTTP requests — invisible in server access logs, and not fixable at any layer below the action. Useforgejo/upload-artifact(anddownload-artifact), whose single patch is to make that check return false.@actions/cache, including viaSwatinem/rust-cache@v2, uses it to select the v1 API:if (isGhes()) return 'v1'. v1 isACTIONS_CACHE_URL+_apis/artifactcache/, which forgejo-runner’s cache proxy implements.
Same check, opposite outcome. No cache found. is the successful-but-empty
branch; an unreachable cache server goes to catch/reportError instead, so
that message alone never means broken plumbing.
Host ephemerality, and why it is not on
Wiping the runner on every boot was considered and rejected:
- It conflicts with the cache.
rust-cacherestores arbitrary files into~/.cargoandtarget/, so a cache preserved across the wipe carries poisoning through it — and a cache not preserved costs the cold column above on every single run. - forgejo-runner already implements GitHub-style PR-scoped cache write
isolation: writes from a pull request go to
refs/pull/N, reads fall back to the shared scope. A PR cannot poison the cache the base branch reads.
So the honest statement is that the host is durable and the job is disposable, and the thing that makes that acceptable is how little the host holds.
TLS, DNS and mail
One certificate
Certificates are security.acme (acme.nix): one certificate named
hu-tao.dev, with every infra.certSubdomains entry as a SAN, issued over
DNS-01 through Cloudflare. Because the certificate is named after the apex
(not after the first domain, as certbot does), reordering the SAN list cannot
silently issue a second lineage under a new name.
- caddy does not manage certificates. It reads the acme directory
read-only; each site names its files explicitly (
tls fullchain.pem key.pem), which turns off caddy’s own management. NixOS names the keykey.pem, not certbot’sprivkey.pem. - DNS-01 only touches
_acme-challengeTXT records, so no name on the certificate needs an A record or a reachable port 80 — which is why the apex itself can be on it, and why the tailnet-only names are ordinary SANs. - The cert directory is group-owned by
caddy(notacme), so the unprivileged caddy container reads it by group membership rather than byCAP_DAC_OVERRIDE, which it drops.reloadServicesrestarts caddy and the mailserver after a renewal. - Every site emits HSTS. Caddy adds nothing of the sort on its own, and
until it was added no vhost here carried it. What it buys is the first
request: every name is https-only and caddy already redirects
http→https, but that redirect is a plaintext round trip an attacker on the path can answer instead — sslstrip againstmail.’s login form, say. Tailnet sites carry it too; the plaintext redirect block deliberately does not, because a browser ignores HSTS on a plaintext response. Nopreload: that is a submission to a list baked into browser binaries, removal takes months, and it would bind the apex and therefore names this caddy does not serve.
The ACME account contact is infra.acmeEmail, and it is ivan@hu-tao.dev —
the mailserver this repo runs. It used to be ivan@hu-tao.org, a Google-hosted
mailbox nothing here describes, so expiry warnings were arriving somewhere this
repo knows nothing about. Changing it registers a new ACME account: lego
keys its account directory by contact address, so the next renewal registers
afresh rather than updating the existing registration. Issuance is unaffected.
Three things follow that option — security.acme.defaults.email, dozzle’s
admin record, and var.caa_iodef in tofu, which is a separate default that has
to be kept equal by hand.
Who may issue — CAA
Without a CAA record, any of the ~150 CAs in the public trust stores may
issue a certificate for hu-tao.dev, and a misissuance is a valid certificate
for this domain in someone else’s hands. DNSSEC does not help: a certificate is
not a DNS answer, so signing the zone says nothing about who may sign for its
name.
The list came from the live certificates, not from acme.nix. Three
issuance paths feed this zone and only one of them is this server — which is
exactly the shape of mistake that makes CAA dangerous, because a record that
omits a real issuer breaks renewals silently, roughly 30 days before an
expiry:
| CA | Who uses it |
|---|---|
letsencrypt.org | lego on the VPS over DNS-01 — and Netlify, which serves the apex (75.2.60.5 / 99.83.231.61) and also issues through Let’s Encrypt |
pki.goog | Cloudflare Universal SSL. www is a proxied CNAME, so Cloudflare terminates TLS at its own edge, currently with Google Trust Services |
Plus an iodef pointing at mailto:ivan@hu-tao.dev — the only way these
records ever report that someone tried.
No issuewild records are published from here, deliberately: RFC 8659 makes
issue govern wildcard issuance when no issuewild is present, and the
tempting hardening — issuewild ";" to forbid wildcards outright — would break
that Cloudflare edge certificate, because it is a wildcard.
Read the result honestly
Cloudflare publishes CAA records of its own when it is the DNS provider, so that Universal SSL stays renewable, and they never appear in a plan. Measured immediately after the apply on 2026-09-20, eight appeared alongside the three tofu owns, and the zone now answers:
issue comodoca.com, digicert.com, letsencrypt.org, pki.goog, ssl.com
issuewild the same five
iodef mailto:ivan@hu-tao.dev
So issuance went from roughly 150 CAs to five, not to one. That is a real
reduction and a modest one, and it is the ceiling for as long as anything in
this zone is proxied — the three records tofu publishes cannot narrow past what
Cloudflare adds back. The only route to the tighter list is to stop needing
Universal SSL at all: un-proxy www, which today is a 301 to an apex Netlify
serves anyway. That is a website decision rather than a security one.
dig CAA hu-tao.dev is the truth; the resource in tofu/ is only the part
tofu owns.
DNSSEC
Signing was switched on in the Cloudflare dashboard long before tofu knew about it, and sat pending because the DS record was never published at the registrar — a signed zone with no chain to the root is an unsigned zone with extra steps. The DS is now published and the chain validates:
dig +dnssec @1.1.1.1 hu-tao.dev SOA | grep -E '^;; flags:.* ad'
dig +dnssec @8.8.8.8 hu-tao.dev SOA | grep -E '^;; flags:.* ad'
tofu output dnssec_ds prints the half a human pastes into the registrar. That
half cannot be automated — Hostinger ships no OpenTofu provider for domain
management — and it does not need to be: a DS is write-once for the life of the
zone.
cloudflare_zone_dnssec.main is adoption only, and its lifecycle block is
the point. An apply that touches the resource puts key material back in play,
and a key rotation after the DS is published takes the whole domain dark for
validating resolvers — mail included. The block is what stops a stray diff from
becoming that action.
The new blind spot worth naming: a broken DNSSEC chain is a whole-domain
outage that the monitoring cannot see. kuma-check probes from the VPS, whose
resolver may not validate, so it would keep reporting green while the rest of
the internet gets SERVFAIL.
Mail deliverability depends on three things agreeing
They are set in three different places, and nothing checks that they match:
flowchart TB
ptr["<b>PTR (rDNS)</b><br/>tofu/rdns.tf<br/><code>smtp.hu-tao.dev</code>"]
dms["<b>DMS hostname</b><br/>mailserver.nix<br/><code>smtp.hu-tao.dev</code>"]
mx["<b>MX target</b><br/>Cloudflare<br/><code>smtp.hu-tao.dev</code>"]
ptr <--> dms
dms <--> mx
mx <--> ptr
The PTR belongs to the primary IP, not the server. That is what makes an IP handover carry mail reputation to a new box with no DNS change — see Migration off Ubuntu.
- SPF is
-alland hard-codes the IP. - DKIM’s public half is published from
tofu/while its private half is a sops secret mounted into DMS. The two must be halves of one key, or every recipient fails the signature. - DMARC is
p=quarantine.
SMTP/IMAP (25/465/587/993) are published directly — an MX must be reachable
at the host, so none of it can sit behind cloudflared. Only the webmail is
proxied, at mail..
postmaster@ and abuse@ are aliases, declared in a read-only
postfix-virtual.cf mounted over the DMS config volume (mailserver.nix).
Before that, postmaster@ answered 550 5.1.1, which also swallowed the box’s
own root mail — and RFC 5321 requires it to be deliverable. setup alias add
now fails on purpose: aliases are a git change, because an alias that exists
only in the dms_config volume is state this repo does not describe.
The Cloudflare tokens
There are two, with different scopes, and confusing them produces a failure
that looks like nothing is wrong: security.acme falls back to a self-signed
certificate and starts its dependent services anyway, so caddy and DMS come up
serving a placeholder. The check is the issuer, not the reachability:
ssh -p 2222 hutao@<host> sudo cat /var/lib/acme/hu-tao.dev/cert.pem \
| openssl x509 -noout -issuer
# want: issuer=C=US, O=Let's Encrypt ...
# bad: issuer=CN=minica root ca ... <- placeholder, DNS-01 failed
Test a token directly rather than by triggering ACME — Let’s Encrypt caps failed validations at 5 per account per hostname per hour, and lego burns one per attempt:
curl -sS -H "Authorization: Bearer $TOKEN" \
'https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev'
The service stack
Containers are virtualisation.oci-containers (docker backend), one module per
service under modules/containers/. Two conventions hold across all of them
(containers/default.nix):
- Data is a named docker volume, never a bind-mounted host directory — so restic covers a new service the moment it declares a volume. See Data and backups.
- Config comes from the Nix store, read-only (0444). A store path changes
with its content, so systemd recreates the container on a config-only change
— which is what the old Ansible setup’s
recreate: alwayswas working around.
Restart policy is forced to always, and logs go to the journal — never
json-file, which would put unbounded logs on the disk and break the forgejo
jail that reads the journal.
What runs, and how it is reached
| Service | Exposure |
|---|---|
caddy | 80/443 (+443/udp) for the public sites, and 8880/8443 for the tailnet-only ones — the second pair is kept private by its absence from the firewall’s allow-lists, and nftables rewrites tailscale0’s 80/443 onto it so those URLs carry no port. Built locally with the caddy-ratelimit module (caddy.withPlugins), not the stock image |
cloudflared | Tunnel connected, but nothing routes through it — git/music/mail/smtp are unproxied A records straight to the VPS, so caddy serves them directly |
mailserver | SMTP/IMAP direct on 25, 465, 587, 993 — an MX must reach the host |
webmail | roundcube, proxied at mail. |
forgejo | SSH on 22, so clone URLs need no port; HTTP via caddy at git.; site admin (/admin, /api/v1/admin) tailnet-only |
navidrome | 127.0.0.1:4533, reached only through caddy at music. |
kuma | proxy network only, reached at status.; the admin socket only on the tailnet copy of that name |
searxng | proxy network only, reached at search.; the only public site behind basic_auth, with caddy rate_limit in front of the bcrypt |
dozzle | dozzle. over the tailnet, and still 8080 directly — the direct port is deliberate, since this is what you open when caddy is the broken part |
grafana | host networking, :3000, tailnet only; also grafana., which caddy reaches at the docker bridge address because host networking is invisible to docker’s DNS |
tempo | host networking, OTLP 4317/4318 bound to 0.0.0.0; kept private by the firewall’s input chain, not by the bind address |
minecraft | 25565; RCON on loopback only (25575). Heap 1 G floor / 6 G ceiling — see below |
minecraft2 | second world, MC 1.21.1 on the java21 image (world 1 is 26.1.2/java25) with its own mod list; 25566, reached via the _minecraft._tcp.mc2 SRV record; RCON on loopback only (25576). Heap 1 G / 4 G |
serenity-bot-0 | nothing published; an outbound Discord gateway client, on the botnet network. tokio-console on 127.0.0.1:6669 |
serenity-redis | botnet only, no published port, no volume — a cache with a Discord fallback |
postgres | not a container — a host service; unix socket + loopback only, never on botnet |
pgbouncer | not a container — a host service; 6432, reachable from botnet and the tailnet, kept private by the firewall’s input chain |
syncthing | not a container — a host service; GUI on 8384, tailnet only; also syncthing., where caddy must rewrite the Host header or syncthing’s rebinding check answers 403 |
endlessh-go | not a container — a host service; the SSH tarpit on 222, 2022 and 22222, public on purpose; its metrics on 127.0.0.1:2112. See Observability |
prometheus | not a container — a host service; loopback-only :9090, scraping the tarpit for grafana |
The Forgejo Actions runner is not on this list any more: it lives on its own machine. See CI runner isolation.
Why the Minecraft heap is two numbers
itzg’s MEMORY sets -Xms and -Xmx to the same value, so the JVM commits
the whole heap at startup and never gives any of it back. That is why an idle
world with nobody on it sat at 4.38 G of RSS at 0–1% CPU. Both worlds therefore
set INIT_MEMORY and MAX_MEMORY separately — a low floor, the same ceiling as
before — rather than MEMORY.
The floor alone is not enough. G1 only uncommits at the end of a GC cycle, and
an idle server triggers no GCs at all, so the heap would stay at its
high-water mark forever. -XX:G1PeriodicGCInterval=300000 (JEP 346) is the
other half: one concurrent cycle per five idle minutes, which is what actually
returns the pages. The two are a pair — either one on its own does nothing
useful.
RSS is still not heap. Metaspace, the code cache, GC structures and direct
buffers live outside -Xmx, so the ceiling is not a bound on what the container
reports.
Forgejo owns port 22, so the host’s sshd is on 2222, and that is the way
in — Tailscale SSH is off. Keeping 22 is what lets git remotes stay portless:
ssh has no service discovery — it reads no SRV record — so anything else has
to be spelled out in every clone URL or every client’s ssh config.
caddy is built here, not pulled
caddy is one of three locally built images (with pages-hook and the serenity
bot, below), and the only one built to add a module. It is built locally with
the caddy-ratelimit module compiled in (caddy.withPlugins, wrapped in a
minimal dockerTools image), because stock caddy has no rate limiting — and
rate limiting is load-bearing for every site, see
Access control.
It is still caddy 2.11.4 — the version tracks nixpkgs, which matches the tag
the official image used — and keeps the full container hardening: non-root uid
from ids.nix, read-only rootfs, --cap-drop=ALL, and tmpfs for /data,
/config and /tmp.
Images are pinned where the tag moves
Most images here are repo:tag on a release tag, which is both a version and a
promise that the content behind it does not change. Three are not, and they
carry a digest as well (forgejo carries one too, on a release tag; the table
lists only the images whose tag itself moves):
| Image | Why the tag alone says nothing |
|---|---|
itzg/minecraft-server:java25 | a rolling JRE tag; a re-pull can swap the JRE under a live world |
itzg/minecraft-server:java21 | the same, and it matters more — world 2’s mods are compiled against Java 21, and mixin/ASM on a newer JDK is the classic crash |
redis:8-alpine | a rolling minor tag |
A digest also makes the image archive’s restore sound: a pinned pull either returns those exact bytes or fails, so falling back to an archived copy cannot silently substitute a different image.
Renovate is configured to match this. Its regex manager captures
currentDigest as an optional group — without it the tag group would
swallow the digest and leave the version unparseable — and the rule that keeps
itzg/minecraft-server off automatic version bumps is split so that
matchUpdateTypes: ["digest"] stays enabled. Before the pin, disabling that
package read as caution and meant the opposite: a rolling tag has no version to
bump, so there was nothing to be deliberate with, and the content moved on the
next pull with no PR ever saying so. A digest PR is that missing signal.
docker.autoPrune is weekly and its flags are empty, so it is a plain
docker system prune — dangling layers only, never a tagged or digest-
referenced image, and never a volume. The runner’s podman prune is the one that
runs --all, and what that costs is covered in
CI runner isolation.
The serenity bot is built during activation
Upstream publishes no image, so serenity-bot.nix builds one from upstream’s
own Dockerfile off a full-SHA fetchgit pin. The build runs in the
activation script, not in its unit: that is the only phase of a switch where
the old container is still serving, so the ~3½-minute compile costs no outage.
stopIfChanged = false on the containers and the builder keeps them up until
then, and serenity-bot-image.service only does the build at boot, when docker
is not up during activation. The tag embeds the rev plus a hash of the enabled
features and rustflags, so a features-only change rebuilds too.
Sharding is two numbers, shards and instances, both 1. The per-instance
ranges are computed and checked exhaustively at eval time up to 16 shards.
Above one instance redis stops being optional: the AI locks and rate limits
are per-process without it. Until then serenity-redis is a cache with a
Discord fallback: in-memory only, no volume, 256 MB LRU.
Deliberately not containers
- postgres + pgbouncer (
postgres.nix) — the first service whose data is not a docker volume, so its backup (pg_dumpall) is not optional; it is the only thing covering that data. On the host so a second service can share it without either owning the other’s volume. postgres is unix-socket and loopback only; pgbouncer (6432) is reachable frombotnetand the tailnet. - syncthing (
syncthing.nix) — a user service with state in~/.config/syncthingand folders under~/syncthing, paths kept byte-identical to the old box because navidrome bind-mounts~/syncthing/Musicand the node’s device ID is derived from the TLS keypair in the config dir. Tailnet-only GUI on 8384. - CoreDNS (
caddy.nix) — answers split DNS for the half-public names, bound to the tailnet address on :53. See split DNS. - Renovate (
renovate.nix) — a daily host timer, on the VPS rather than the CI runner so its cross-org write token never reaches an untrusted machine. - The Actions runner, which is now a different machine entirely.
Access control
Every service has exactly one gate, and they differ by what the service itself supports.
| Service | Gate |
|---|---|
| grafana | real login from sops; anonymous-Admin off, sign-up off |
| dozzle | bcrypt hash from sops, in a users.yml copied to a stable path |
| minecraft RCON | password from sops, loopback only; one password per world (25575 / 25576) |
| serenity bot | discord token + AI key + db password, all sops |
| kuma | no seeding mechanism — the first visitor creates the admin account and the route then closes. Create it immediately after the first deploy. |
| syncthing | GUI password from sops via guiPasswordFile, declared rather than inherited |
| pages-hook | HMAC of the Forgejo system webhook; the secret is in sops and must match the hook’s Secret field |
| searxng | no accounts at all — caddy’s basic_auth is the entire access control; the bcrypt hash is a sops secret handed to caddy via an env file |
Site administration is tailnet-only
Where an app’s administration lives on paths its public side never uses, caddy
answers those paths with a 404 on the public listener and serves the same name
again on the tailnet listener with them open (blockedPaths in
modules/containers/caddy.nix). Tailnet devices resolve those names to the
tailnet address through
split DNS, so the admin side
works from any of them with nothing configured per device.
- kuma —
/socket.io/is only the admin login and dashboard; the public status page loads everything from/api/status-page/*. kuma’s login is a message inside that socket, so no HTTP rate limit can count attempts at it: not exposing the socket is the only real second layer. - Forgejo —
/admin(the site admin panel) and/api/v1/admin/*(its API: users, orgs, runner tokens, cron, system webhooks). Personal and org settings stay public. A stolen password or session can still use what the account owns, but not administer the instance from outside the tailnet. Forgejo also routes//adminand/api//v1/admin/*to the same handlers; caddy’s path matcher merges slashes, and a local replay confirmed those 404 too.
The block is its own handle, not a bare respond: caddy orders handle
before respond, and the git site’s upstream sits inside handle blocks for
the runner restriction, so a bare respond there would never run.
Rate limiting, every site by default
Every caddy site is rate-limited unless it sets rateLimit = null, and none
does. It used to be opt-in, and only search. had opted in, which left the
logins on music., mail. and git. open to guessing at whatever speed each
app allowed — Forgejo allows any. Up to three zones per site, all keyed on the
client IP (modules/containers/caddy.nix):
- Site-wide — a flood brake: 300 a minute by default;
git.600;music.,grafana.andsyncthing.1200;search.120. Private ranges, this host’s public address and the CI runners are exempt, so kuma, renovate, pages-pull and CI are never throttled. - Login — 10 POSTs a minute to the login request, with no exemptions:
Roundcube
/?_task=login, Forgejo/user/login(and two-factor, forgot-password, sign-up), Navidrome/auth/login. Nothing on this host posts to a login form, so a private source here can only be a masked client, and a shared bucket failing closed is the right answer. - Basic auth —
git.only: 30 a minute for requests carryingAuthorization: Basic, which is how a password reaches Forgejo over git or the API without ever touching/user/login. Exempt like the site-wide zone, because renovate pushes with its token as a basic-auth password.
These are one layer, not the only one. Roundcube also locks an account after 3 failures a minute; Navidrome 0.64.1 throttles failed Subsonic logins itself; searxng’s bcrypt sits behind the limit. Forgejo has no throttle of its own, so two-factor auth on the account is its second layer.
The mail protocols (465, 587, 993) never pass through caddy. Their only
brute-force control is docker-mailserver’s own fail2ban: six failures in a week
buy a week’s ban. Its exemption list is widened to every docker network by a
mounted fail2ban-jail.cf, because Roundcube’s IMAP logins arrive from a
docker address — without it, six wrong webmail passwords would ban Roundcube
itself and take webmail down for everyone.
searxng, the interesting one
It has no concept of a user, so authentication is the proxy’s job — and that turns out to have a second-order problem.
flowchart TB
c(("client")) --> rl["caddy rate_limit<br/>per client IP"]
rl -- "over limit" --> r429["429<br/>before any bcrypt"]
rl -- "under limit" --> ba["basic_auth<br/>cost-14 bcrypt"]
ba -- "bad" --> r401["401"]
ba -- "good" --> sx["searxng"]
sx -- "image_proxy<br/>many thumbnails per page" --> rl
classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
class r429,r401 bad
- caddy’s
basic_authon thesearch.site is the only thing between the instance and the internet. The credential is passed through caddy’s own{$VAR}env substitution from a sops-rendered env file — not a Nix string, because the Caddyfile is a world-readable store path and a bcrypt hash there is one anyone with a shell could crack. - basic_auth runs a cost-14 bcrypt on every request, and searxng’s
image_proxypulls many thumbnails per results page through caddy — so a password flood could turn bcrypt into CPU exhaustion. caddy’srate_limit(the compiled-in module) caps hits per client IP and returns 429 before the bcrypt runs, orderedbefore basic_auth.
The choice of an in-process limiter over a fail2ban jail is deliberate, and it is the same lesson as the fail2ban blast radius: a misconfigured rate limiter throttles requests, it cannot take the box down.
The env file is read by docker at container start, so it re-resolves the sops generation symlink each time. A mounted template would pin a stale inode — see Secrets.
Identity model
modules/ids.nix assigns uids/gids to services that run under their own
account, as base (1_000_000) + offset, from a hand-maintained append-only
table.
The base clears every allocator the host uses — system users, nixbld,
DynamicUser, subuid blocks — so nothing this repo assigns can ever collide
with something NixOS allocated. The ceiling (2097151) is the largest uid a
ustar/tar header can hold, which matters because these uids end up in restic
snapshots.
The numbers are assigned, not hashed from the service name. A hash becomes immutable the moment the first file is written, and a rename then silently orphans every file the old name owned.
Read config.infra.serviceId.<name>, never a literal — so grep -rn serviceId
finds every use.
Two services have ids: caddy (1) and pages-hook (2), which must own its
bind-mounted /run directory without being root. A group per id is created so
security.acme can chown the cert directory to caddy rather than to acme,
which is what lets the container read its certificate by group membership
instead of by a capability it drops.
The runner’s users are separate
modules/runner/users.nix is its own file rather than a shared import. The
runner holds no age key, so it has no sops-provisioned accounts; its
forgejo-runner user is a plain static system user, deliberately not
DynamicUser.
The reason is ordering: the uuid+secret pair has to be written by a unit that runs before the daemon starts, and a dynamic uid does not exist until the unit it belongs to starts. There would be no stable owner to chown the composed config to ahead of time. See CI runner isolation.
Data and backups
flowchart TB
vols[("docker volumes<br/>/var/lib/docker/volumes")]
pg["postgres (host service)"]
reg["container images"]
dumps[("pg_dumpall<br/>/var/backup/postgresql")]
imgs[("docker save<br/>/var/lib/image-archive")]
mc[("minecraft_data<br/>minecraft2_data")]
daily["restic · daily 00:00–01:00<br/>keep 7d / 4w / 6m"]
weekly["restic · weekly<br/>servers STOPPED"]
b2[("Backblaze B2")]
pg -- "23:15, before restic" --> dumps
reg -- "23:30, before restic" --> imgs
dumps --> daily
imgs --> daily
vols --> daily
mc -- "excluded from daily" --> weekly
daily --> b2
weekly --> b2
Wholesale, not per service
services.restic backs up /var/lib/docker/volumes wholesale, so a
service added later is covered the moment it declares a volume. A backup that
has to be told about each new service eventually stops covering one. The same
rule shapes the other two paths: both are generated from what the configuration
already declares, never from a hand-maintained list.
postgres runs on the host, so its data is not under docker/volumes;
services.postgresqlBackup writes a pg_dumpall (every database plus globals)
there at 23:15, deliberately before restic’s window, so the archived dump
is never up to 23 hours stale.
The image archive
The third path is not data — it is the bytes of every container image, and it exists because a registry reference is not a guarantee that the bytes are still there.
This estate has already been bitten. forgejo 16.0.2 is unpullable: the tag still resolves on codeberg, but a platform manifest inside the index was deleted, so a rebuild from scratch cannot reach it. That was a release tag, not a rolling one, which is the part that matters — pinning a digest does not cause this and staying on a tag does not prevent it. The only thing that helps is owning a copy.
So modules/image-archive.nix runs docker save over every pullable image
at 23:30, into a directory the daily restic job already ships to B2. 23:30
is the same reasoning as postgresqlBackup at 23:15: an archive written after
the backup window reaches B2 a day late.
Locally built images are excluded, because a docker pull can never satisfy
one and each already has a unit that produces it — caddy and pages-hook
(imageFile, built by dockerTools) and serenity-bot-*. serenity-redis is
not in that set: it runs a registry image and is covered like everything
else.
Restoring is automatic
Each pullable container’s unit gained an ExecStartPre that resolves its image
in four steps, and only reaches the last two when the ones above have failed:
already present → docker pull → the archive on disk → restic restore
Digest pinning is what makes that fallback sound rather than merely convenient: a pinned pull either returns those exact bytes or fails, so “pull, else restore” can never quietly substitute a different image. With a floating tag it could. See the service stack.
The restore names --path /var/lib/image-archive rather than taking a bare
latest, because this repository also holds the weekly Minecraft job’s
snapshots and those carry no archive at all.
Not compressed, deliberately
docker save already emits the layer blobs the way the registry stores them,
so zstd measured 38 M → 38 M on redis. Worse, one compressed stream per
archive would defeat restic’s content-defined deduplication — which is what
makes a nightly copy of 2.6 G nearly free and lets the two Minecraft images
share their common base layers.
The Minecraft worlds are a separate job
They are snapshotted with the servers stopped — a live world holds region
files open and is not consistent on disk. That job stops the containers in
backupPrepareCommand and restarts them from backupCleanupCommand (an
ExecStopPost), so the servers return whether restic succeeded or not.
Both worlds share one job, and therefore one downtime window: a second job would mean a second stop/start cycle and a second restic run against the same repository.
Every world volume must be in this job’s
pathsand in the daily job’sexclude. A volume missing from the exclude list is archived hot by the daily run, which is exactly the corruption this job exists to prevent.
The password is the key
The restic password is the encryption key: lose it and every snapshot is unrecoverable. It, and the other things that deliberately live outside this repo, are catalogued in the README’s “Not in this repo” section.
What the runner holds
Nothing that is backed up, and that is the point. The runner’s state is a nix store and an Actions cache — both reconstructible, both worthless to an attacker, neither in any restic job. A runner is replaced, not restored. See CI runner isolation.
Secrets
sops + age (secrets.nix). The age private key lives at
/var/lib/sops-nix/key.txt on the host and is staged there before first
boot by nix run .#install — without it, activation cannot decrypt anything
and the machine boots with no credentials, its own login included.
Two kinds of consumer:
sops.secrets.*— a decrypted file at/run/secrets/…, for things read as a file: the restic password, the DKIM key, the database password, the syncthing GUI password. The last one isowner-ed tohutaorather than left at the default0400 root, becausesyncthing-initreads it asservices.syncthing.userand not as root.sops.templates.*— a rendered file mixing secrets with literal text, for things that wantKEY=valueor a whole config file: most env files here (grep -rn sops.templates moduleslists all sixteen), the pgbouncer userlist, dozzle’susers.yml— and the tailnet auth key, which is the interesting one.
Every key in secrets.nix must exist in secrets.yaml, or
sops-install-secrets fails during activation. This is validated at build
time, so a missing key fails nix build rather than only the deploy — which is
why a new secret is added to secrets.yaml before the module that reads it.
acme_email is deliberately not a secret: security.acme needs it at
evaluation time, and a registration contact address is not a credential. It is
infra.acmeEmail.
The tailnet key is assembled, not stored
tailscale_oauth_client_secret is an OAuth client secret, not a
tskey-auth- key. Tailscale accepts one in place of an auth key and it does
not expire, where the auth keys it replaces capped out at 90 days — leaving the
box one forgotten rotation away from being unable to rejoin its own tailnet
after a rebuild.
sops holds that string and nothing else. The two query parameters that go with
it live in modules/services.nix, in the clear, on purpose:
sops.templates."tailscale-authkey".content =
"${config.sops.placeholder.tailscale_oauth_client_secret}?ephemeral=false&preauthorized=true";
ephemeral=false is mandatory and load-bearing. An OAuth-minted key defaults
to ephemeral=true, and an ephemeral node is removed from the tailnet when it
goes offline — so at the default this VPS would delete itself on every reboot
and rejoin as a new node with a new address, silently invalidating three DNS
records and the policy file’s vps host. Inside ciphertext that is a
one-character mistake nobody can review; in a template it is a line in a diff.
See The tailnet policy.
The stale-symlink trap
A rendered template’s real path is under a generation directory, and .path
only symlinks to it. Docker resolves a symlink at mount time and holds that
inode forever, so a rotated secret never reaches a container that mounts
the template.
flowchart TB
subgraph broken["✗ mounted template"]
t1["/run/secrets/rendered/x<br/>(symlink)"] --> g1["…/generation-4/x"]
d1["docker mount"] -. "resolved once,<br/>inode pinned" .-> g1
g2["…/generation-5/x<br/>(rotated)"]
t1 -. "now points here" .-> g2
d1 -.->|"never sees it"| g2
end
broken ~~~ ok
subgraph ok["✓ two fixes"]
f1["copy to a stable path first<br/><code>dozzle-users.service</code><br/><code>mailserver-dkim.service</code>"]
f2["pass as an env file<br/>docker re-reads at container start<br/>caddy · cloudflared"]
end
Both fixes work because they re-resolve the symlink at container start rather
than pinning it at mount. The env-file form is the lighter of the two and is
why searxng’s bcrypt hash reaches caddy that way — see
Access control.
The runner has no secrets at all
mkRunner does not pass sops-nix.nixosModules.sops. The runner holds no age
key and can decrypt nothing in secrets.yaml; the only credential on the box
is its own Forgejo registration pair, delivered through Hetzner user-data.
The instance-id check does not make that pair useless elsewhere, and it is
worth being precise about what it does do. The stamp lives in the staged file
at /var/lib/forgejo-runner-identity/userdata and is compared against the live
metadata value, so a pair staged for one box is not silently adopted by
another. Forgejo knows nothing about Hetzner instance IDs — it accepts the
uuid and secret from anywhere. Anyone holding that pair can register a second
runner daemon against the instance and start receiving jobs.
Which is why modules/runner/firewall.nix drops link-local traffic from job
containers: the same user-data that delivers the pair is readable with one
curl from inside a job unless the forward chain stops it. See the metadata
service is the host’s alone.
This is asserted, not merely intended: checks.runner-has-no-secrets is a
nix build, because a check that is only evaluated never runs its builder.
Never decrypt to inspect
Reading secrets.yaml to “check” something is how secrets end up in a
terminal, a log, or a pull request. The manifest in secrets.nix is the list
of what exists; the live box is the place to verify that a secret arrived
(systemctl status, the consuming service’s own health), not the ciphertext.
Observability
flowchart TB
bot["serenity bot"] -- OTLP --> tempo["tempo :4317/4318<br/>host net"]
tempo --> graf["grafana :3000<br/>tailnet only"]
graf ~~~ cont
cont["every container"] -- journald --> doz["dozzle :8080<br/>tailnet"]
cont -- journald --> jctl["journalctl"]
tarpit["endlessh-go<br/>:222 · :2022 · :22222"] -- "metrics :2112" --> prom["prometheus :9090<br/>loopback"]
prom --> graf
doz ~~~ kuma
kuma["kuma<br/>proxy net"] --> cad["caddy"] --> status["status.hu-tao.dev"]
check["kuma-check<br/>every 5 min"] -- "probes the PUBLIC page" --> status
check -- ping --> hc(("healthchecks.io"))
classDef alert fill:#8c5a1f,stroke:#4d330d,color:#fff
class check,hc alert
tempo, grafana and prometheus use host networking and are kept private by the input chain (Network), not by their bind address. kuma publishes no port and is reached only through caddy.
The self-hosted status page paradox
kuma cannot report its own host being down. That is closed by kuma-check, an
out-of-band timer that probes the public status page every five minutes and
pings healthchecks.io. Healthchecks alerts on the absence of a ping, so
silence becomes the alert rather than an all-clear.
The principle is worth naming once: a monitor that reports nothing when it cannot run is not a monitor.
The bot’s dashboards
Four of the dashboards seeded from modules/containers/grafana-dashboards/
into the Provisioned folder are the bot’s: overview, guild, user and DMs,
linked so a guild or user row drills into its own view with the time range
carried across. Every panel on them is TraceQL against tempo. The fifth is the
tarpit’s. Any other dashboard lives only in grafana_data, which
restic already covers.
UI edits save (allowUiUpdates), but the directory is one store path, so
editing any file in it re-seeds all five and discards UI edits across the
folder. Export a dashboard back into git before touching its neighbours.
Two query guards are bound to a date:
- a second target on
span.guild_id =~ "Some.*"keeps spans from before 2026-09-17, whenguild_idwas recorded in its Debug spelling; span.attachment_urls != nilon the span tables, becauseattachmentsandlinkschanged from string to integer on 2026-09-18, and the tempo plugin panics (HTTP 500, “No data”) on a column whose type changes mid-result.
Delete both once tempo’s retention no longer reaches those dates.
The SSH tarpit
modules/tarpit.nix runs endlessh-go on 222, 2022 and 22222, three common
alternative SSH ports; the dashboard splits by port. It accepts the connection
and sends a random line a second, forever, so a client waiting for the SSH
version string waits for hours. It guards nothing (sshd on 2222 is key-only
either way); it is there to waste bots’ time and to count them.
Its Endlessh dashboard, in the Provisioned folder, shows connections,
time trapped, and a map of where they came from. The map comes from
-geoip_supplier=ip-api: each new client address is looked up at ip-api.com
over plain http, so bot addresses go to that third party. Its free tier allows
45 lookups a minute; a client beyond that is still trapped, just without a
point on the map.
prometheus exists for this dashboard alone, on loopback :9090, with grafana’s second datasource pointed at it. Its data is a 15-day window of bot statistics and is not backed up.
A tarpit port is open in three places: modules/tarpit.nix, the input chain in
modules/firewall.nix, and tofu/modules/hetzner-firewall. The last one only
takes effect on tofu apply.
Where to look when something is wrong
| Symptom | First place |
|---|---|
| a site is down | dozzle at :8080 directly — it is what you open when caddy is the broken part |
| a container will not start | journalctl -u docker-<name> |
| a deploy failed | the deploy output itself; then journalctl -u <unit> for the unit it named |
| mail is not delivered | the issuer check in TLS, DNS and mail, then DMS’s own log |
| a CI job never starts | journalctl -u forgejo-runner on the runner, reached with ssh -J hutao@vps:2222 |
| traces are missing | tempo is on host networking; check the input chain admits the bot’s bridge |
Deploying
Two paths here, in order of how often you’ll walk them: redeploy (constantly) and first deploy to an existing NixOS machine (once per machine). Installing onto hardware that has no NixOS at all is Provisioning.
There are two deploy targets, and one of them is only reachable through the other:
deploy .#vps # the VPS
deploy .#runner-forgejo-runner # the CI runner, via ProxyJump through the VPS
Everything below assumes the ssh key is loaded, because it is passphrase- protected and nothing here can prompt for it:
ssh-agent -a /tmp/hutao-agent.sock >/dev/null 2>&1
SSH_AUTH_SOCK=/tmp/hutao-agent.sock ssh-add ~/.ssh/id_ed25519
export SSH_AUTH_SOCK=/tmp/hutao-agent.sock
A Permission denied (publickey) from any command here almost always means the
agent is gone, not that a key is missing on a server.
1. Redeploy — the everyday path
nix flake check # optional; deploy builds anyway
deploy .#vps
Without deploy-rs — same result, no automatic rollback:
nix develop # provides nixos-rebuild on a non-NixOS workstation
nixos-rebuild switch --flake .#vps-hetzner --target-host hutao@vps --use-remote-sudo
If deploy-rs itself ever becomes unavailable, note that an unresolvable flake
input stops the flake evaluating at all — so that fallback needs the input
removed from flake.nix first, not just a different command.
That is the whole thing. deploy-rs builds locally, pushes the closure,
activates it, then waits for a fresh connection to confirm the box is still
reachable. If it cannot reconnect, the machine rolls itself back to the
previous generation without being asked.
What it protects and what it does not:
| Failure | Caught by |
|---|---|
| firewall / sshd / networking change locks you out | deploy-rs auto-rollback |
| unbootable kernel or initrd | GRUB generation menu, 5s timeout at boot |
| a container fails to start | not auto-rolled back — see below |
The last row is deliberate. deploy-rs confirms reachability, not service health. A crashlooping container is visible and you still have ssh, so:
ssh -p 2222 hutao@vps 'systemctl --failed; systemctl status docker-<name>'
ssh -p 2222 hutao@vps 'sudo nixos-rebuild switch --rollback'
Rolling the whole system back because one container is unhappy is usually the wrong reflex — fix it forward.
sequenceDiagram
participant W as workstation
participant H as host
Note over W: build the closure locally
W->>H: push closure, activate
W--xH: close the connection
W->>H: reconnect, FRESH connection
alt reachable within confirmTimeout
W->>H: confirm
Note over H: new generation kept
else unreachable
Note over H: rolls itself back, unattended
end
confirmTimeout is 120 s — long enough for every container to be recreated on
a config change, short enough that a hung activation is not an outage. The
activation timeout is 900 s, raised from 300 s for the serenity-bot image
build, which runs in the activation script while the old container is still
serving; serenity-bot-image.service is only the boot-time fallback. Measured on
this host on 2026-09-04, cold cache including the base image pulls: 3m28s.
The runner is deployed through the VPS
The runner is deliberately off the tailnet, and inbound ssh is narrowed to the
VPS /32 at both layers — its cloud firewall and its own nftables — so the
only route in is a jump:
deploy .#runner-forgejo-runner # sshOpts carry ProxyJump=hutao@vps:2222
ssh -J hutao@vps:2222 root@46.225.61.172 # by hand
Magic rollback matters more here than anywhere: a mistake in
modules/runner/firewall.nix locks out the only path to the box, and the jump
host cannot help with that. Deploy the VPS first and the runner last when a
change touches both — the runner’s route in is defined by the VPS’s outbound
rules, so a VPS deploy that has not landed yet means a runner you cannot
reach.
The jump hop goes to 2222. ProxyJump=hutao@vps:2222 names the port on
purpose: a bare vps lands on 22, which on the tailnet used to be Tailscale SSH
and is forgejo’s now that Tailscale SSH is off. Both deploys use the VPS’s own
sshd with an ordinary key, which is why 2222 is listed in the policy and called
deploy-critical there.
Ports and names, so nothing surprises you
vpsis the tailnet node name, which is whydeploy.nodes.vps.hostnameis a name and not an address. It is notnetworking.hostName(hu-tao): renaming the node in the Tailscale console breaks deploys. It survives the primary-IP handover during a migration, so the same command works before and after cutover.- ssh is on 2222. Port 22 belongs to forgejo, so that git clone URLs need no port. Tailscale SSH is off, so 2222 with a normal key is the only ssh in.
If a deploy fails with “lacks a signature by a trusted key”
nix.settings.trusted-users must include @wheel (it does, in
modules/nix.nix). If you ever deploy to a machine that predates that setting,
you cannot push to it — build on the box instead:
rsync -a --delete --exclude .git -e 'ssh -p 2222' ./ hutao@vps:nixos-image/
ssh -p 2222 hutao@vps 'cd nixos-image && sudo nixos-rebuild switch --flake .#vps-hetzner'
That is also the bootstrap for the very first deploy after an install.
2. First deploy to a machine that already runs NixOS
Same as a redeploy, with two one-time steps:
# 1. Trust the host key, or deploy-rs fails with "Host key verification failed"
# and no way to answer the prompt.
ssh-keyscan -p 2222 -H hu-tao >> ~/.ssh/known_hosts
# 2. Confirm the box can decrypt its own secrets before relying on it.
ssh -p 2222 hutao@vps 'sudo ls /run/secrets/ | wc -l' # expect 31
31 is 30 secrets plus the rendered/ directory; the root and user password
hashes live in /run/secrets-for-users. Recompute it when a module adds a key:
nix eval --json .#nixosConfigurations.vps-hetzner.config.sops.secrets \
--apply 's: builtins.length (builtins.filter (v: !v.neededForUsers) (builtins.attrValues s)) + 1'
Then deploy .#vps.
What makes this automatic (and what used to break it)
Every item below is now in the config. They are listed because each one, when missing, produces a machine that installs with no error and then does not work — the worst failure shape there is.
| Setting | Where | Without it |
|---|---|---|
boot.initrd.availableKernelModules with virtio | modules/hardware.nix | NixOS’s default set is bare-metal only. The initrd cannot see /dev/sda, root never mounts, and the box sits in an emergency shell while the provider still reports it running. |
| GRUB, not systemd-boot | modules/boot.nix | Hetzner Cloud boots legacy BIOS — there is no /sys/firmware/efi. systemd-boot installs cleanly and leaves an unbootable machine. |
efiInstallAsRemovable, canTouchEfiVariables = false | modules/boot.nix | There is no efivarfs in BIOS mode; bootloader installation fails outright if it tries to write NVRAM. |
time.timeZone | modules/boot.nix | Unset means NixOS does not manage /etc/localtime, so docker creates a directory there and every container that bind-mounts it dies with “not a directory”. |
nix.settings.trusted-users = @wheel | modules/nix.nix | deploy-rs cannot push: “lacks a signature by a trusted key”. |
| ssh on 2222 | modules/services.nix | Port 22 is forgejo’s. Also needs a matching rule in the Hetzner edge firewall, which is separate from the host’s nftables. |
The VM test cannot catch any of these. nixos-anywhere --flake .#vps --vm-test validates disko, GRUB and that the system boots — genuinely useful,
and it is what proved GRUB-on-BIOS works. But the NixOS test harness injects its
own virtio modules and its own networking, so a config that boots in the test can
still be unbootable on real hardware. Treat a passing VM test as “the layout and
bootloader are sane”, never as “this will boot on the server”.
Verifying a machine is actually healthy
Not “the deploy said success” — these:
ssh -p 2222 hutao@vps '
systemctl is-system-running # want: running
systemctl --failed # want: empty
sudo ls /run/secrets | wc -l # want: 31
sudo docker ps --format "{{.Names}} {{.Status}}"
for u in caddy forgejo mailserver webmail kuma navidrome searxng pages-hook minecraft minecraft2 \
grafana tempo dozzle cloudflared serenity-bot-0 serenity-redis; do
echo "$u restarts=$(systemctl show -p NRestarts --value docker-$u)"
done
systemctl is-active postgresql pgbouncer serenity-bot-image'
Non-zero NRestarts means a crashloop that systemctl is-active will happily
report as active, because systemd restarts it fast enough to look healthy.
And confirm the certificate is real rather than the self-signed placeholder that
security.acme installs when issuance fails — services start either way, so
nothing looks wrong until you check the issuer:
ssh -p 2222 hutao@vps 'sudo cat /var/lib/acme/hu-tao.dev/cert.pem' \
| openssl x509 -noout -issuer -enddate
# want: issuer=C=US, O=Let's Encrypt, ...
# bad: issuer=CN=minica root ca ... <- placeholder, DNS-01 failed
Before triggering ACME, test the Cloudflare token directly — Let’s Encrypt caps failed validations at 5 per hour and lego spends one per attempt:
ssh -p 2222 hutao@vps 'sudo bash -c "
T=\$(cat /run/secrets/cloudflare_api_token)
curl -sS -H \"Authorization: Bearer \$T\" \
https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev"'
An empty result array with success: true means the token cannot see the zone
— which reads as success if you only check .success.
If deploy-rs ever disappears from GitHub
The risk is bigger than losing deploy: flake inputs are fetched from source,
not from the binary cache, so an unresolvable input means the flake stops
evaluating entirely and nixos-rebuild --flake fails too.
Three mitigations, in order of effort:
-
Nothing breaks while the store path is present. A locked input already realised in
/nix/storeis not refetched.nix flake archivecopies every input into the store on purpose, andnix-store --gcis what would remove them again. -
Keep a copy:
nix flake archive --to file:///path/to/mirrorwrites all inputs somewhere you control. -
Cut it out — a three-part edit to
flake.nix, after which option 2 above is the only deploy path:- delete the
deploy-rsentry frominputs - delete
deploy-rsfrom theoutputs = { ... }argument list - delete the
deployandchecksoutputs, anddeploy-rs.packages.${system}.defaultfrom the devShell
Nothing in
modules/references it, so the machine configuration itself is unaffected. - delete the
Provisioning a machine
Installing NixOS onto hardware that has none, and the maintenance that only comes up around an install: restoring data into a fresh service, and growing the disk after a resize.
For updating a machine that already runs NixOS, see Deploying.
Bare metal — a brand new server
This is meant to be close to one command. It is, provided the machine’s quirks are already in the config — see “What makes this automatic” below.
cd tofu
cp terraform.tfvars.example terraform.tfvars && $EDITOR terraform.tfvars
tofu init
tofu plan # READ IT. Abort on any "destroy and then create" of hcloud_server.
tofu apply
tofu creates the server with your ssh key attached at creation. It does not
install NixOS — that is deliberate. Run the install separately:
nix run .#install -- root@$(tofu -chdir=tofu output -raw vps_ipv4)
tofu used to own the install through nixos-anywhere’s module. That was removed:
the module declares a null_resource whose creation runs a full install, so
any plan made without it already in state — a fresh clone, a lost state file, a
state rm — quietly proposes reinstalling a running mail server. Infrastructure
and OS installation are now separate on purpose.
Installing onto a server that already exists
One command:
nix run .#install -- root@<ip>
It verifies the age key decrypts secrets.yaml before starting, stages it
into a temporary extra-files tree at 0600, and selects vps-hetzner. Extra
arguments are passed through to nixos-anywhere (--debug, --build-on-remote).
Equivalent by hand, if you ever need to vary it:
mkdir -p /tmp/extra/var/lib/sops-nix
install -m 0600 ~/.sops-nix/key.txt /tmp/extra/var/lib/sops-nix/key.txt
nix run github:nix-community/nixos-anywhere -- \
--flake .#vps-hetzner \
--target-host root@<ip> \
--extra-files /tmp/extra
--extra-files is not optional. Without the age key at
/var/lib/sops-nix/key.txt, sops-install-secrets fails during activation and
the machine boots with no credentials at all — including its own root and user
passwords. Check the key decrypts before installing:
SOPS_AGE_KEY_FILE=~/.sops-nix/key.txt sops -d --extract '["email"]["postmaster"]' secrets.yaml
If the target only accepts a key you do not hold, Hetzner rescue mode is the way
in — enable_rescue accepts an ssh_keys list, unlike rebuild, which
re-injects whatever was attached at creation:
curl -X POST -H "Authorization: Bearer $HCLOUD_TOKEN" -H 'Content-Type: application/json' \
-d '{"type":"linux64","ssh_keys":[<key-id>]}' \
https://api.hetzner.cloud/v1/servers/<id>/actions/enable_rescue
curl -X POST -H "Authorization: Bearer $HCLOUD_TOKEN" \
https://api.hetzner.cloud/v1/servers/<id>/actions/reset
Rescue is a normal Linux with the disk unmounted, which is exactly what nixos-anywhere wants.
Which configuration to install
| Attr | Disk | Use |
|---|---|---|
.#vps | /dev/vda | the local QEMU VM (nix run .#default) |
.#vps-hetzner | /dev/sda | anything on Hetzner Cloud |
They are the same closure; only the disk device differs, and boot.loader.grub.device
is derived from disko so the two can never disagree. Installing .#vps on
Hetzner fails at disko because /dev/vda does not exist there.
Restoring a postgres dump into a new service
services.postgresqlBackup writes a pg_dumpall to
/var/backup/postgresql/all.sql.zstd nightly, and restic carries it — so the
usual restore is one command:
zstd -d < /var/backup/postgresql/all.sql.zstd | sudo -u postgres psql
Seeding a service from a dump made elsewhere is different, and the ordering
gets one shot. The bot runs its sqlx migrations automatically on startup, so if
it reaches an empty database first, its migrations create the schema and the
dump’s CREATE TABLEs then collide with it. Restore before the container’s first
start.
Read the dump before running anything — two of its properties decide the commands, and guessing either one wrong fails halfway through:
head -40 ~/serenity-bot-db.sql # pg_dump (needs a target db) or pg_dumpall (has its own CREATE DATABASE)?
grep -m5 'OWNER TO' ~/serenity-bot-db.sql # which role does it expect to exist?
A plain-SQL dump emits ALTER TABLE … OWNER TO <role>, which hard-fails under
ON_ERROR_STOP=1 if that role is absent. The cheap fix is to make the config
match the dump — role and db at the top of modules/postgres.nix — rather
than to rewrite the dump.
Then, for a pg_dump of a single database:
sudo systemctl stop 'docker-serenity-bot-*'
sudo -u postgres psql -c 'DROP DATABASE IF EXISTS serenity_bot;'
sudo -u postgres psql -c 'CREATE DATABASE serenity_bot OWNER serenity;'
sudo -u postgres psql -v ON_ERROR_STOP=1 -d serenity_bot -f ~/serenity-bot-db.sql
sudo systemctl start docker-serenity-bot-0
ON_ERROR_STOP=1 is not optional: without it psql reports success after
skipping every statement it could not apply, which leaves a half-populated
database that looks restored.
Confirm the bot treats the schema as current rather than migrating it:
journalctl -u docker-serenity-bot-0 -n 50
sudo -u postgres psql -d serenity_bot -c 'table _sqlx_migrations order by version desc limit 5;'
Growing the disk after a Hetzner resize
Resizing the volume in the Hetzner console changes the block device and nothing
else. disk-config.nix declares root as size = "100%", so a fresh install
fills the new disk correctly — but disko only partitions at install time, so a
running machine needs this once, by hand.
The symptom is that the space is invisible rather than merely unused:
lsblk -o NAME,SIZE # sda 152.6G, but sda3 only 75.3G
sfdisk --list-free /dev/sda # "Unpartitioned space: 0 B" <- lying
0 B is the tell. GPT keeps a backup header at the end of the disk, so
after a resize the table still describes the old geometry — last-lba points at
the old final sector, and within that table the last partition genuinely does
fill the disk. sfdisk says so out loud if you read past the numbers:
GPT PMBR size mismatch (160006143 != 320004095) will be corrected by write.
The backup GPT table is not on the end of the device.
Nothing can see the free space until that header moves. Neither sgdisk nor
growpart is in the system closure; pull them from the pinned nixpkgs rather
than adding them permanently for a once-per-machine job:
nix build --no-link --print-out-paths 'nixpkgs#gptfdisk^out' # sgdisk
nix build --no-link --print-out-paths 'nixpkgs#cloud-utils' # growpart
nix copy --to ssh://hutao@vps --no-check-sigs <both paths>
Take a backup first — restic runs daily, so force one if anything since the
last run matters, and save the partition table as the rollback artifact:
sudo systemctl start postgresqlBackup # get the databases into the dump
sudo systemctl start restic-backups-b2 # then offsite
sfdisk -d /dev/sda | sudo tee /root/sda-parttable-$(date +%F).sfdisk
Then four steps, in this order:
sudo sgdisk -e /dev/sda # 1. move the backup GPT to the true end of disk
sudo growpart /dev/sda 3 # 2. extend the LAST partition into the new space
sudo partx -u /dev/sda # 3. make the running kernel re-read the table
sudo resize2fs /dev/sda3 # 4. grow ext4 online; safe while mounted
Step 1 is the one that is easy to skip and impossible to work around: without
it, step 2 finds no free space. Step 2 only changes the partition’s end
offset, so no data moves — this works because root is the last partition. Step 4
needs the resize_inode feature, which is present.
Restore path if step 1 or 2 goes wrong: sfdisk /dev/sda < /root/sda-parttable-*.sfdisk,
from Hetzner rescue mode if the box will not boot.
e2fsck will look alarming afterwards, and it is lying
sudo e2fsck -fn /dev/sda3
# ... ********** WARNING: Filesystem still has errors **********
This is expected and is not evidence of damage. e2fsck cannot meaningfully
check a mounted read-write filesystem: the kernel holds allocation state that
has not reached the on-disk bitmaps, and -n skips journal recovery. So it
reports bitmap differences, wrong free counts, and “deleted inode has zero
dtime” for files unlinked while still open. Declining every fix under -n is
what sets the error banner.
Prove it rather than worry about it — run it twice and compare:
sudo e2fsck -fn /dev/sda3 | grep "count wrong ("
sleep 20
sudo e2fsck -fn /dev/sda3 | grep "count wrong ("
Different numbers each run means it is tracking live writes. Identical numbers
would be the thing to investigate. What actually matters is
tune2fs -l /dev/sda3 | grep "Filesystem state" reporting clean.
A real check needs the filesystem offline: add fsck.mode=force fsck.repair=yes
to the kernel command line for one boot, or run it from rescue mode.
Why not LVM
It would not have helped much here. Growing into new space would still need
the GPT header relocated and the partition extended before pvresize,
lvextend, resize2fs — three commands instead of two, for the same outcome.
Where it would pay is snapshots before a risky migration, and reallocating space
between volumes. Switching means a reinstall, since disko partitions only at
install time, so it belongs to the next bare-metal build rather than to a
resize.
CI
Two workflows, one per forge. Forgejo reads .forgejo/workflows and falls
back to .github/workflows only when that directory is absent — a fallback,
not a union — so the presence of .forgejo/workflows/ci.yml is what keeps
Forgejo off the GitHub file. Delete it and Forgejo silently starts running a
workflow written for GitHub, which is how this repo once ended up with a red
run on git.hu-tao.dev.
| File | Runs on | Jobs |
|---|---|---|
.forgejo/workflows/ci.yml | the dedicated runner box | one: check |
.forgejo/workflows/pages.yml | the dedicated runner box | one: pages — builds this book, main only |
.github/workflows/ci.yml | the GitHub mirror | two: lint and evaluate |
| Check | What |
|---|---|
| lint | pre-commit run --all-files, then gitleaks across the full history |
| evaluate | evaluates all three nixosConfigurations, then nix flake check --no-build, then builds deploy-schema and its guard, runner-firewall-ordering, runner-has-no-secrets, pages-pull-strips-symlinks, and the minecraft-mods package |
The mirror is push-only. Commit here and let it flow across; anything edited on GitHub is overwritten by the next sync, and CI can lag a push until Forgejo’s mirror job runs (Synchronize Now in the repo’s mirror settings).
Why the two files differ
Not tidiness — each difference is forced.
-
Job layout. The Forgejo file is one job; the mirror’s is two. The runner keeps a nix store between runs now that it has its own box, but the
.#cishell still has to be realised, and two jobs would pay that in parallel oncapacity: 2. -
Actions, or none at all. The Forgejo job runs on the
nixlabel —nixos/nix, which already contains Nix — so there is nocachix/install-nix-actionto run and no Nix to download per run. That image carries no node, and every JavaScript action is executed by a node binary inside the job container, so the CI file has nouses:whatsoever and does its owngit fetchin place ofactions/checkout.That is a property of this label, not of the runner. Since 2026-09-21 there is also
nix-node— nix and node in one locally built image — on whichuses:works normally. These two workflows have not moved to it and do not need to; see the job image for when a repo should.pages.ymlis the exception that proves it. It needs one action —upload-artifact, which has no shell equivalent — so it puts the dev shell’s node on$GITHUB_PATHfirst, which is whynodejsis in thecishell despite nothing here being a node project. Skipping that step fails the job withcrun: executable file 'node' not found in $PATHbefore the action runs at all.skavexandhutao/compresspublish the same way.The image is not only a saving, it is the fix: on
ubuntu-latest(node:22-bookworm)install-nix-actionexits 127, because the branch it takes without systemd runssudo mkdir -p /etc/nixand that image has no sudo. The job is already root, so the sudo bought nothing to begin with.
permissions: is a GitHub-only field. Forgejo ignores it with a workflow
warning, which is why the Forgejo file omits it rather than carrying a line
that does nothing.
Evaluation, not a build
Both forges evaluate rather than build. It catches what actually breaks this repo — a typo’d option, a missing module argument, an infinite recursion — without asking a runner to realise a multi-gigabyte closure.
--no-build is load-bearing. deploy-rs’s deploy-activate check references
the system closure, so a plain nix flake check builds the whole system, and
since deploy-rs follows our nixpkgs its binary is a cache miss and is
compiled from source.
Five checks and one package are built rather than evaluated, and each for a reason:
| Check | Why it must build |
|---|---|
deploy-schema | validates deploy.json against deploy-rs’s schema — the config is otherwise only exercised by a real deploy. deploy-schema-rejects-bad-input is its guard: feed the validator a node with no hostname and fail if that is accepted, so a validator that silently reads nothing cannot pass forever |
runner-firewall-ordering | greps the evaluated nftables ruleset for rule order. A check that is only evaluated never runs its builder, so its failure branch would be inert |
runner-has-no-secrets | same reason — it catches the runner growing a sops-install-secrets unit, which a copy-pasted module import would do |
pages-pull-strips-symlinks | asserts the symlink strip and chmod sit between unzip and copy, then runs those exact lines against a hostile zip — evaluating it would never run the builder |
minecraft-mods | a package, not a check: it is __noChroot because it queries Modrinth, and in checks it would fail nix flake check on any sandboxed machine. The GitHub job passes --option sandbox false; the Forgejo image already has sandbox = false |
runner-firewall — the real two-node VM test with real packets — is not in
CI. It needs /dev/kvm, and these are shared-vCPU Hetzner instances with no
nested virtualisation; qemu’s TCG software emulation was measured at roughly 5×
slower just to boot two minimal nodes. It stays hand-run on a machine with KVM:
nix build .#checks.x86_64-linux.runner-firewall -L
Writing a workflow that runs here
The runner is a different machine from the VPS, with no access to its volumes and a four-path allow-list to its Forgejo instance. Four consequences:
- The label decides whether
uses:works at all.nixhas no node, so a JavaScript action cannot execute on it — the failure iscrun: executable file 'node' not found in $PATH, before the action runs. Picknix-nodefor a workflow that wantsactions/checkoutor any other action, and pick it for any private repo: a hand-writtengit fetchis anonymous unless the author threads the job token through it, which works on a public repo and fails on a private one. Full list of labels and what each carries: the job image. upload-artifactmust be the Forgejo fork. GitHub’s bundles@actions/artifactv2, which decides a Forgejo instance is GitHub Enterprise Server and throws before opening a socket — zero HTTP requests, invisible in access logs. Useforgejo/upload-artifact@v5; a bareuses:resolves againsthttps://code.forgejo.org, which is correct for it.- Publishing writes an artifact, never a volume. See Pages.
- A bare
uses:resolves againstDEFAULT_ACTIONS_URL— which defaults tohttps://data.forgejo.org, a mirror ofactions/*and nothing third-party. A third-party action must name its host, or it fails withremote: Not found. PointingDEFAULT_ACTIONS_URLat github.com would fix it instance-wide, at the cost of making every bareuses:resolve to whoever holds that name on an open-registration forge.
Caching works normally — see
why the cache works when artifacts did not.
No cache found. is the successful-but-empty branch, not a broken cache.
Known gaps
- Actions are pinned by moving tag rather than commit SHA.
pages.ymlstill puts the dev shell’s node on$GITHUB_PATHbefore its oneuses:step. Moving that job tonix-nodewould delete the step; it has not been done, because the same file isskavex’s and the two are kept identical deliberately.- The
nix-nodeimage is rebuilt only whenflake.lockmoves, so a CVE in its node or curl waits for a flake bump. - The artifact pull has no size cap (
--max-timebounds time, not bytes). - Job containers do not yet use
--userns=auto.
Pages
pages.hu-tao.dev is a static site caddy serves out of one docker volume.
The URL layout is the directory layout, with no rewriting anywhere:
<pages_data>/<owner>/<repo>/index.html → https://pages.hu-tao.dev/<owner>/<repo>/
This book is published that way, at
https://pages.hu-tao.dev/hutao/vps/docs/.
The direction reversed
A publishing job used to mount the pages_data volume and write into it —
that is what the old in-container runner’s one-entry valid_volumes allow-list
was for. A runner on its own box cannot do that, and must not: it would be the
runner reaching into the VPS, which is the one thing the split forbids.
So the job uploads an artifact and the VPS fetches it.
flowchart TB
subgraph R["runner box"]
job["workflow job"]
end
job -- "upload artifact<br/>named <b>pages</b>" --> fj
subgraph V["hu-tao"]
direction TB
fj["Forgejo"]
fj -- "POST action_run_success<br/>over the docker bridge" --> hook["pages-hook<br/>verify HMAC, touch a file"]
hook -- "systemd .path" --> pull["pages-pull"]
clock(["hourly timer<br/>safety net"]) --> pull
pull -- "discover repos,<br/>fetch what changed" --> fj
pull --> vol[("pages_data")]
vol -- read-only --> caddy["caddy"]
end
caddy --> url(["pages.hu-tao.dev/<owner>/<repo>/"])
classDef net fill:#2d4a7c,stroke:#16233c,color:#fff
class hook net
An hourly timer starts pages-pull as well, and that is a safety net rather
than the mechanism — see When it runs.
Every connection is initiated on the VPS, and in fact never leaves the host — the artifact is in Forgejo’s own storage, in a container on the same box.
The workflow runs on main only. An earlier version built on pull requests
and gated the upload step with if: push && main instead, which meant the
riskiest step was skipped in every rehearsal: the branch introducing the
workflow went green having never once run it, and the failure landed on main
at merge. Gate the workflow, not the step.
No credential. The publishing repos are public and Forgejo serves
/api/v1/repos/<owner>/<repo>/actions/artifacts anonymously (verified
2026-09-19 against the live instance: 200 with a bare JSON array body). The day
a private repo publishes, this needs a sops token with read:repository — and
not before. Do not add one speculatively.
Not a gh-pages branch, which would be the idiomatic shape. That needs
git-receive-pack, which caddy now denies to runner addresses: the
allow-list carries info/refs and git-upload-pack and nothing else, so a
runner can clone and cannot push. Artifacts ride the twirp ArtifactService
instead, sidestepping the need entirely. See
CI runner isolation.
Adding a repo
One thing: give the repo a workflow that uploads an artifact named exactly
pages. That is the whole of it. Nothing is added to this repo, and no deploy
is run.
pages-pull enumerates every repo on the instance through
/api/v1/repos/search and publishes any that holds a live artifact by that
name. Uploading it is the opt-in, the same way enabling Pages is a repo-level
act rather than something the hosting provider does for you.
It was not always so. There used to be an infra.pagesRepos list in
modules/options.nix, so adding a page meant a commit here and a
deploy .#vps — a rebuild of the machine that serves mail, in order to publish
a static site. The list existed because of a misreading of the runner split:
the note said the pull side “has to be told what to look for”, when it only has
to be told how to find out.
Two things fall out of the change:
- A repo that has never published is not a special case. It has no
pagesartifact, so it is not discovered, and nothing is logged. Under the list this was a repo you had named in error and worth a line in the journal; now it is simply most of the instance. - A renamed or deleted repo stops being discovered. The old list had to be
edited by hand when that happened, or the unit failed every five minutes
forever —
hutao/critical-forestwas carried as a comment explaining exactly that. The served tree is still left in place, so caddy keeps answering that path until someone removes it.
An empty discovery is treated as an error, not as “no repos publish”. This instance always has repos, so zero means the search endpoint moved or started refusing us — and the damage would be every published site silently freezing at its current content while the unit kept exiting 0.
The artifact is untrusted content
It was built on the CI runner, the machine this
estate treats as hostile. Layers 1–5 over there are address-and-port-and-path
controls; an artifact is the one thing that crosses the boundary carrying
content, and the caddy allow-list has to admit the ArtifactService route or
publishing does not work at all.
So pages-pull reduces the unpacked tree to what a published site is actually
made of, before anything is pointed at it:
unzip -q "$tmp/pages.zip" -d "$tmp/out"
find "$tmp/out" ! -type f ! -type d -delete
chmod -R a-s,go-w "$tmp/out"
Info-ZIP already refuses the two obvious escapes — it strips ../, it warns
stripped absolute path spec from /x, and it will not write through a
symlink. What it does do is restore one, target and all, and the cp -a
that follows preserves it. That is enough on its own, because caddy’s
file_server follows a symlink out of its root, and modules/acme.nix
group-owns the certificate directory by caddy so that caddy can read
key.pem. One ln -s /etc/caddy/certs/key.pem x in a published artifact would
otherwise serve the apex certificate’s private key — and every one of its
eleven SANs — at https://pages.<domain>/<owner>/<repo>/x.
A published site is files and directories; a symlink in one has no legitimate
use here. They are deleted rather than rejected so that one malformed
artifact cannot wedge a repo’s publishing, and the match is on the entry itself
and never its target, so this cannot follow a link out of $tmp.
The filter is an allow-list of entry types, not a list of known-bad ones.
A zip cannot carry a device node or a socket, but it can carry a fifo, and
cp -a would preserve it into the served tree — where caddy’s file_server
blocks forever on open(2) the first time anyone requests that path. Deleting
everything that is not a regular file or a directory also means the next entry
type an archive format learns does not get a free pass.
The chmod covers the other thing a zip carries: a mode. Info-ZIP already
drops setuid and setgid on extraction, so a-s is belt-and-braces against an
extractor that someday does not — but go-w is not redundant. A world-writable
file in the pages volume is one that any process on this host can rewrite after
publication, with caddy serving whatever it finds on the next request. It runs
after the delete, so there is no symlink left for chmod -R to follow out
of $tmp.
checks.pages-pull-strips-symlinks holds both: it asserts the strip sits
between the unzip and the copy and that the chmod sits between the strip and
the copy — a chmod ordered before the strip would walk a tree that still
contains symlinks and follow one out, reintroducing the exact escape the strip
exists to close. It then evals those exact lines — lifted out of the
evaluated unit, not retyped — against a tree unpacked from a hostile zip, with
a fifo and a world-writable file planted on top of it.
Why the unit is still root
pages-pull runs as root, and the reason is the destination rather than the
work. The tree lands in /var/lib/docker/volumes/<vol>/_data, and
/var/lib/docker is 0710 root:root — nothing unprivileged can even traverse
it. Adding the unit’s user to the docker group is not the alternative: that
group is root-equivalent by design. Dropping the privilege properly means
moving pages off a docker volume onto a plain bind-mounted directory, which is
a volume migration rather than an edit.
So root is bounded instead. The unit parses a network-fetched ZIP with
unzip, which is the one place in this module where hostile content meets a
parser, and the mitigation that matters is that winning there reaches nothing
worth having. Under ProtectSystem=strict with ReadWritePaths naming only
/var/lib/docker/volumes, plus PrivateTmp, NoNewPrivileges, an empty
CapabilityBoundingSet and SystemCallFilter=@system-service, a compromised
unzip can write to the pages volume and its own private /tmp and nothing
else — not /var/lib/sops-nix, not /etc, not /home, not the docker socket,
with no way to gain a capability back.
ReadWritePaths names the volumes parent, not the volume’s own _data
path, on purpose: _data does not exist until docker first creates the volume,
and a ReadWritePaths entry that does not resolve fails the unit on a fresh
box.
When it runs
A Forgejo system webhook on action_run_success, so publishing is an event
rather than a poll. The whole path is:
| trigger | one system hook in Site Administration, firing for every repo on the instance |
| target | http://<dockerBridgeGateway>:<pagesHookPort>/hooks/pages-pull |
| auth | HMAC-SHA256 over the body, read from X-Hub-Signature-256 |
| receiver | modules/pages-hook.nix — webhook(1) in a container on the proxy network, no published port |
| effect | touches one file; a systemd .path unit starts pages-pull as root |
It never leaves the proxy network. Both Forgejo and the receiver are containers on it, so the delivery is container-to-container. There is no caddy site, no published port, no public listener — and no firewall rule at all, because the port never exists on the host.
That is also the correction to an earlier claim here: this used to say a webhook cost “an HTTP receiver on the mail server, which is not a trade worth making”. It assumed the receiver had to be public.
A system hook, not a per-repo hook. Per-repo would reintroduce exactly the per-repo setup step that discovery removed. One hook covers everything, including repos that do not exist yet.
The receiver parses nothing. Any successful Action Run pokes pages-pull,
which is idempotent and cheap. Reading the payload would trade a slightly
smaller number of no-op runs for a coupling to Forgejo’s ActionPayload
schema.
The listener holds no privilege. It runs as its own uid from
modules/ids.nix, with a read-only rootfs and every capability dropped, and
its entire capability is touching one file in a bind-mounted directory; the
.path unit does the privileged half. A network-facing process that can run
systemctl start is a network-facing process that is root-adjacent.
The hourly timer stays, and is now a safety net rather than the mechanism. A webhook is a delivery and deliveries are lost — the receiver can be down mid-deploy, Forgejo’s retries can run out, the hook can be switched off in a web form nothing here can see. Each of those leaves a site frozen with no error anywhere. The sweep makes the worst case “stale for up to an hour”.
Forgejo has to be told the destination is allowed
ALLOWED_HOST_LIST defaults to external, which permits public addresses and
blocks private ones — so out of the box Forgejo refuses to deliver and the
webhook silently never fires. modules/containers/forgejo.nix sets it to
pages-hook.
The container name, not an address, and that is the security-relevant part.
The list matches hosts, not host:port. While the receiver ran on the host,
this had to name 172.17.0.1 — which also permitted a webhook aimed at
anything else bound there, and grafana (3000), syncthing’s GUI (8384), tempo
(4317/4318) and pgbouncer (6432) all bind 0.0.0.0 and answer on it. Forgejo
webhooks can use GET and record the response body in their delivery
history, so that was a read primitive with an exfiltration channel: whoever
could create a webhook could read tailnet-only services without being on the
tailnet.
A name works because the matcher is MatchHostName(host) || MatchIPAddr(ip) —
a name pattern alone is sufficient, so no private address needs allowing and
172.17.0.1 stops matching at all. One destination, nothing else.
The failure mode is worth knowing because everything on the RECEIVING side looks correct while it happens: the socket is listening, the rule matches, and there is no dropped packet, no connection refused and nothing in the receiver’s journal — because nothing is ever sent.
Look at the sender instead. Forgejo logs the refusal as an error on its own service:
journalctl -u docker-forgejo | grep -i 'unable to deliver webhook'
and the hook’s settings page shows the same thing under Recent deliveries.
The one hand-kept value
The hook’s Target URL and secret live in a web form, so nothing in this repo
can verify they match infra.pagesHookPort and
forgejo/system_webhooks/pages_pull/secret. A mismatch
is at least loud in two places: a 403 in journalctl -u pages-hook, and a
failed delivery in the hook’s own history in Site Administration.
Failure isolation
Each repo runs in its own subshell, so one bad repo — renamed, deleted, made private, a network blip, a corrupt zip — cannot take down the rest of the loop. A transient failure heals itself on the next tick; a persistent one is the failure worth guarding against, because left unguarded it would permanently block every repo listed after it, and silently: the unit “succeeding” on the repos before the broken one looks no different from everything being fine.
So the loop logs it, keeps going, and fails the unit at the end. failed is
set, not incremented — whether it was one repo or all of them, the answer
is “look at the journal”, not a count.
journalctl -u pages-pull --since -1h
systemctl list-timers pages-pull
Checking it worked
curl -sSI https://pages.hu-tao.dev/hutao/vps/docs/ | head -1
Runbook
Commands, in the order you are likely to want them. Everything here runs on the
VPS over ssh on port 2222 (ssh -p 2222 hutao@vps); the runner’s equivalents
are at the bottom, and they are reached differently.
Is anything broken
systemctl --failed
systemctl list-units 'docker-*'
journalctl -u docker-forgejo -f # every container logs to the journal
Everything else
systemctl list-units 'docker-*'
journalctl -u docker-forgejo -f # every container logs to the journal
systemctl status acme-renew-hu-tao.dev.timer # renews well before expiry; a no-op most days
systemctl start acme-hu-tao.dev.service # force a renewal check
systemctl status restic-backups-b2.timer restic-backups-minecraft.timer
restic-b2 snapshots # wrapper with the repo and password wired in
systemctl status postgresqlBackup.timer # 23:15, deliberately BEFORE restic's 00:00-01:00 window
systemctl start postgresqlBackup # dump every database now
psql -h 127.0.0.1 -p 6432 -U serenity serenity_bot # through pgbouncer, from the tailnet
fail2ban-client status forgejo-ssh
nft list table inet f2b-table # where the bans actually are
nft list table inet nixos-fw
tailscale status # peers, and this node's own address
tailscale whois 100.109.115.12 # THIS node: `Tags: tag:vps` or the ACL does not apply
systemctl status image-archive.timer # 23:30, ahead of restic's 00:00-01:00 window
systemctl start image-archive # docker save every pullable image now
ls -la /var/lib/image-archive # one .tar + one .id per image
When an image cannot be pulled
Nothing to do — every pullable container’s unit runs ensure-image before it
starts, which tries the local image, then a pull, then the nightly archive,
then restic. journalctl -u docker-<name> names whichever step it reached.
The recovery path can be exercised by hand without touching a running service:
restic-b2 restore latest --path /var/lib/image-archive \
--include /var/lib/image-archive/<slug>.tar --target /tmp/rt
--path is not optional: this repository also holds the weekly minecraft
snapshots, and a bare latest can name one of those, which carries no archive.
--target / is what ensure-image uses, because restic recreates the absolute
path under the target. The slug is the image reference with /, : and @
each replaced by _. Verified end to end on 2026-09-21.
The weekly minecraft job stops both worlds’ servers, snapshots, and starts
them again from ExecStopPost — so they come back whether restic succeeded or
not. A live world is not consistent on disk: the server holds region files open
and writes them in place, which is why the daily backup excludes that volume and
this job exists.
On the runner
The runner is off the tailnet on purpose, so every command goes through the VPS:
ssh -J hutao@vps:2222 root@46.225.61.172
systemctl status forgejo-runner
journalctl -u forgejo-runner -f # a job that never starts shows here
journalctl -u forgejo-runner-identity # the uuid/secret compose step
podman ps # job containers, one per running job
podman images # localhost/forgejo-ci-nix-node must be here
systemctl restart forgejo-runner-ci-image # reload it if the prune ate it
du -sh /var/lib/forgejo-runner/cache # the Actions cache
nft list table inet nixos-fw # the one-way rules; counters included
nft list counters is the quick check that the one-way rule is doing something:
vps_allowed_out should climb while jobs run, and vps_blocked_out should stay
where it was. See CI runner isolation.
The ingress pair answers the other direction — ssh_from_vps climbs every time
you open the jump above, and ssh_blocked is anyone else trying:
ssh -J hutao@vps:2222 root@46.225.61.172 nft list counter inet nixos-fw ssh_from_vps
ssh -J hutao@vps:2222 root@46.225.61.172 nft list counter inet nixos-fw ssh_blocked
Renovate
Renovate runs on the VPS, not the CI runner: its token can write across
hutao/* and skavex/*, and a runner treated as hostile never holds it.
systemctl status renovate.timer # daily 12:00 UTC, ±15 min
systemctl start renovate # run now
journalctl -u renovate -n 200
Nothing opens until it is ticked on the Dependency Dashboard
(dependencyDashboardApproval in renovate.json5). With the CVE scanner
dropped, reading that issue is the only path for security updates.
Failure modes and recovery
The design is shaped by which failures roll back automatically and which do not. This page is the table of both, and then the three that are recent scars.
| Failure | Caught by | Recovery |
|---|---|---|
| firewall / sshd / networking change locks you out | deploy-rs auto-rollback | automatic |
| unbootable kernel / initrd | GRUB generation menu (5 s at boot) | pick the previous generation |
| a container fails to start | nothing — deploy confirms reachability, not health | nixos-rebuild --rollback or fix forward |
nftables reload wiped docker’s chains | nothing automatic; the symptom is the next container start failing | systemctl restart docker, and keep flushRuleset = false |
| a rotated sops secret didn’t reach a container | nothing — the container holds a stale inode | copy-to-stable-path or env-file, see Secrets |
| tofu plan shows an unexpected diff on unchanged infra | you, reading the plan | the state is wrong, not the infra — never apply; refresh/import, verify on the box |
| a fail2ban jail bans a docker/bridge address | nothing — a forward-chain reject on an internal IP downs every container | systemctl stop fail2ban; keep private ranges in ignoreIP |
| a deploy restarts every container at once and one racy unit exits non-zero | deploy-rs aborts — and its deactivation stops every container | see the abort that was worse than the failure |
tofu apply fails with test(s) failed (400) on the tailnet policy | the policy’s own tests, run by Tailscale at apply time | the narrowing is not live yet — see the bootstrap deadlock |
| a container image is gone from the registry | nothing — the pull fails at the next start, on an image that has worked for months | automatic: ensure-image falls back to the nightly archive, then to restic — see the image archive |
| the DNSSEC chain breaks | nothing here — kuma-check probes from the VPS, whose resolver may not validate | dig +dnssec @1.1.1.1 hu-tao.dev SOA from off-box; until then the domain is dark for validating resolvers |
The fail2ban blast radius
A forward-chain fail2ban jail has a blast radius the size of the whole box. If it ever bans an internal source it rejects all forwarded traffic, not one attacker — every container goes dark at once.
That is why the sites are rate-limited in caddy rather than banned in nftables: an in-process limiter can only throttle, it cannot take the forward plane down. See Access control.
The host’s ignoreIP includes 172.16.0.0/12, which holds every docker network
here, so no jail can ban a bridge address at all — structurally, not because a
regex happens not to match one. docker-mailserver’s own fail2ban carries the
same exemption, for the same reason inside its own namespace.
Before adding any forward-chain jail, ask what happens when it bans
172.30.0.x.
The abort that was worse than the failure
The least intuitive row, and the widest. A failed activation does not leave the box on the previous generation running. It leaves it on the previous generation’s configuration, with nothing started.
flowchart TB
u["nix flake update<br/>new nixpkgs"] --> r["every unit changes<br/>→ every container restarts at once"]
r --> race["one racy unit exits non-zero<br/>(3-second transient)"]
race --> abort["deploy-rs samples unit state,<br/>sees one failed unit, ABORTS"]
abort --> deact["deactivation stops<br/><b>every</b> container"]
deact --> out["15 healthy services down, mail included<br/>6 minutes"]
deact --> lock["switch-to-configuration refuses:<br/>Could not acquire lock"]
classDef bad fill:#8c2f2f,stroke:#4d1a1a,color:#fff
class abort,deact,out,lock bad
On 2026-09-16 a three-second transient in one non-critical container therefore cost fifteen healthy services, mail included, for six minutes. The abort was more destructive than the failure it was responding to.
Recovery, in order:
systemctl restart docker
# then start the container units by hand — switch-to-configuration will refuse
# with "Could not acquire lock" while deploy-rs still holds it
The structural fix is to order fragile units behind a readiness gate rather than letting them race the restart. Full writeup: 2026-09-16 — flake update rollback.
tofu state recovery
State is gitignored — it holds every value tofu has ever read. A clone has
none, and apply from no state builds a second server and moves DNS to it.
tofu/imports.tf prevents that declaratively: an import block per live
resource, inert while state tracks it, active when it does not. tofu init && tofu plan then rebuilds state.
A correct recovery plan reads 32 to import, 0 to add, 1 to change, 0 to destroy — the one change being three provider-side booleans on
hcloud_server.vps that the importer never sets. Anything else means stop.
Expect a second change, tailscale_acl.main, until the policy in this repo
has been applied at least once: the import reads whatever the tailnet is
currently serving, and the diff against tailscale-policy.hujson is the point.
Read it before applying — that resource replaces the whole document, so
anything the live policy carries and the file does not is deleted with nothing
in the plan naming it as a loss. See
The tailnet policy.
Three resources are in state that nothing declares — hcloud_network,
hcloud_network_subnet and hcloud_server_network.vps, left over from the
private network that was deleted. They cannot be given import blocks
(Configuration for import target does not exist), and plans refresh them and
then omit them from resource_changes entirely, which is recorded in
tofu/imports.tf as not understood rather than explained away. Clearing them
means tofu state rm.
The sharp edge (learned the hard way, 2026-09-05): the hcloud provider never reads
public_netinto state, so a post-import plan proposes adding it — and on this resource that detaches the primary IPs before reattaching. Applying it once took the mail IP off a running host.server.tfnow carrieslifecycle.ignore_changes = [public_net], so the block can never become an action. The primary IP is its own resource with delete protection precisely so a mistake here is minutes of downtime rather than a lost address.
A plan that disagrees with what you know to be true is a state problem, not
an infrastructure problem. Never reach for apply to make it agree.
Runner-specific failures
| Symptom | Cause | Fix |
|---|---|---|
| jobs queue and never start | the runner cannot reach Forgejo, or its identity is wrong | ssh -J hutao@vps:2222 root@<runner>, then journalctl -u forgejo-runner |
403 not permitted from a CI runner in a job | the workflow hit a path outside the four-path allow-list | widen the allow-list deliberately, or change the workflow — see runner isolation |
upload-artifact fails with no HTTP request at all | GitHub’s action, not the Forgejo fork | forgejo/upload-artifact@v5 |
| you cannot ssh to the runner | its route in is defined by the VPS’s outbound rules | deploy the VPS first; the runner has no tailnet and no other path |
every job on nix-node fails in ~3 s, on unchanged workflows | the daily podman system prune --all deleted the locally built job image | systemctl restart forgejo-runner-ci-image — or just deploy; both now reload it. See the garbage collector eats it |
A runner that is broken past recovery is replaced, not repaired — it holds a nix store and a job cache and nothing else. Rebuild in place, never destroy and recreate: CX instance types are limited availability.
Development
nix develop # or `nix develop -c zsh`
The shell carries everything: nixfmt, statix, sops, age, nixos-anywhere,
nixos-rebuild, opentofu, deploy-rs, pre-commit, gitleaks, markdownlint-cli2 and
mdbook. .#ci is a deliberately smaller subset — see CI.
Running it locally
Tests:
nix run github:nix-community/nixos-anywhere -- --flake .#vps --vm-test
The actual system (infinitely more useful):
QEMU_OPTS="-vnc :0" nix run .#default
And ssh into it from another terminal — ssh -p 2222 hutao@127.0.0.1 (there’s
no place like 127.0.0.1). Port 2222 on both ends, because sshd moved off 22 so
forgejo could publish it.
The VM disk is 32G (virtualisation.vmVariantWithDisko). The disko default of 2G
leaves ~987M for / once the ESP takes its gigabyte, which cannot hold the sixteen
declared images — the VM fills up mid-boot and every service that needs disk fails
in a way that reads like a bug in that service. This applies to nix run . only;
the Hetzner disk is sized by the provider.
Working from a Mac
The devShell is built for all four mainstream systems — x86_64-linux,
aarch64-linux, aarch64-darwin, x86_64-darwin. Every tool in it,
nixos-anywhere, nixos-rebuild and deploy-rs included, exists on each.
Secrets, formatting, statix and the hooks work unchanged.
What does not carry over is building the system closure. The outputs that
describe the box — nixosConfigurations, packages, apps — are
x86_64-linux only, so nix build ., nix run . (the QEMU VM) and
nix run .#install fail on anything else: a Mac has no Linux builder at all,
and an aarch64-linux workstation is the wrong architecture. Deploys work, but
only if the build happens somewhere else:
deploy -s --remote-build .#vps # build on the VPS itself
nixos-rebuild switch --flake .#vps-hetzner \
--target-host hutao@vps --build-host hutao@vps --use-remote-sudo
-s is not optional here. deploy-rs runs nix flake check first, and every
check in this flake reaches nixosConfigurations.*.system.build.toplevel, so
the check itself is an x86_64-linux build:
error: build of '…-10-acme.conf.drv^*' failed: platform mismatch
Required system: 'x86_64-linux' Current system: 'aarch64-darwin'
There is nothing to keep by skipping selectively — checks.aarch64-darwin
exists but both entries depend on the same Linux closure, so none of them
build here either.
--remote-build then evaluates locally (which darwin does fine), copies the
.drv with nix copy --to ssh-ng://hutao@vps --derivation, and realises it on
the box. That copy needs the ssh user to be a trusted nix user; hutao is in
wheel and modules/nix.nix trusts @wheel, so it already is.
The alternative is a Linux remote builder in /etc/nix/machines (or
nix-darwin’s nix.linux-builder), after which the plain commands above work
as written — including nix run .#install, which is otherwise Linux-only and so
still the reason a first install is done from a Linux machine.
Formatting
Formatting is nixfmt, not nixpkgs-fmt — every .nix file here conforms to
it and the two disagree on multi-argument lambdas, so the wrong one reformats the
whole tree.
nixfmt $(git ls-files '*.nix') && statix check .
tofu -chdir=tofu fmt -recursive && tofu -chdir=tofu validate
Hooks
nix develop -c pre-commit install # once per clone
nix develop -c pre-commit run --all-files
.pre-commit-config.yaml is the single definition — the local commit hook and CI
run the same file, so a check cannot pass here and fail there. It covers nixfmt,
statix, tofu fmt, markdownlint-cli2, the usual whitespace/YAML/merge-conflict
hooks, and gitleaks over the staged diff.
See below for what the markdown hook does and does not do.
gitleaks scans the staged diff rather than the working directory on purpose:
gitleaks dir reads gitignored files, and tofu/terraform.tfvars legitimately
holds live tokens — scanning it would fail the hook forever over a file git will
never accept. History scanning is a CI step instead.
Writing documentation
This book is docs/. mdbook and mdbook-mermaid are in the dev shell:
nix run .#docs # install assets, then serve
nix develop -c mdbook build docs # what CI does
nix run .#docs exists because the build has a prerequisite that is easy to
forget: mdbook-mermaid install docs writes mermaid.min.js and
mermaid-init.js next to book.toml, and book.toml references them. Those
two files are gitignored — 2.6 MB of vendored minified JS whose version is
already pinned by flake.lock — so a fresh clone does not have them and
mdbook build fails until they are written. The app does both steps; the
workflow does the same two commands explicitly.
create-missing = false in book.toml, so a link to a page that does not
exist fails the build rather than publishing a 404.
Diagrams are mermaid, and that is the point
Every diagram here is a ```mermaid fence in the markdown. Forgejo bundles
mermaid 11.16.1, so the same source renders in three places with no build
step: this book, the Forgejo web UI when browsing docs/src/, and a pull
request that changes one.
That is the whole reason they are not SVGs, D2, or the ASCII art they replaced. A picture that only exists after a build is a picture nobody sees while reviewing the change that invalidates it.
Markdown is linted, not formatted, by the hook
markdownlint-cli2 reports a code fence with no language and a paragraph past
80 columns. It rewrites nothing.
Formatting is prettier’s, run by an editor on save rather than by a hook, and
.markdownlint-cli2.yaml is tuned so prettier’s output passes untouched. The
config lists every rule that is off and why. Prettier is deliberately not in
the dev shell: nixpkgs-26.05 carries 3.8.3, which mangles a paragraph
containing both an intraword underscore and an emphasis span.
docs/superpowers/ is not linted — those are agent-generated plan and spec
records, kept as history.
Checking the Minecraft mod lists
nix run .#minecraft-mod-check
It checks every declared Modrinth slug against the exact Minecraft version and
loader each world runs, read from the evaluated config rather than a copy of
it. CI builds the same thing as .#minecraft-mods, which needs
--option sandbox false because it uses the network.
CX33 → CX33 (→CX43) migration
Status: completed. This is the plan as it was run, kept as the record. The old box (
137766340) has since been deleted, so the rollback in step 5 no longer exists andlegacy_server_idsintofu/variables.tfis empty. The serenity bot, which this plan left on the old box, was ported afterwards in8bba9bb. For a future move, reuse the method (the data mapping, populate volumes before first start, the primary-IP handover), not the numbers.
Moving the stack from ubuntu-4gb-fsn1-2 (server 137766340, Ubuntu + Ansible)
to ubuntu-8gb-fsn1-1 (server 163906050, NixOS from this flake). Both in
fsn1, which is the fact the whole plan rests on: primary IPs are
location-bound, so 167.233.24.58 and its smtp.hu-tao.dev PTR can move
between these two machines. Nothing in DNS changes, SPF keeps naming the same
address, and sending reputation carries over intact.
Rescale to CX43 came after the migration. The box is CX43 today.
Data mapping
The old box bind-mounts host directories; this config uses named docker volumes. That makes the transfer a mapping, not an rsync of one tree. ~27 GB total.
| Source (old box) | Size | Destination (new box) | Entries |
|---|---|---|---|
/srv/minecraft/data/ | 15 G | volume minecraft_data | 8107 |
~/syncthing/ | 11 G | ~/syncthing/ (host path) | — |
/srv/forgejo/data/ | 271 M | volume forgejo_data | 1151 |
~/data/navidrome/ | 123 M | volume navidrome_data | 5020 |
/srv/mailserver/data/dms/mail-state/ | 95 M | volume dms_state | — |
volume grafana-data | 50 M | volume grafana_data | — |
volume tempo-data | 16 M | volume tempo_data | — |
/srv/kuma/data/ | 8.7 M | volume kuma_data | 9 |
/srv/mailserver/data/dms/mail-data/ | 2.5 M | volume dms_mail | — |
/srv/mailserver/data/roundcube/db/ | 1.3 M | volume roundcube_db | — |
/srv/mailserver/data/dms/config/ | 12 K | volume dms_config | 1 account |
~/.config/syncthing/ | small | ~/.config/syncthing/ (host path) | — |
Read these sizes as root. Measured as hutao, /srv/mailserver/data reports
15 M; as root it is 110 M, because 24 paths under /srv are owned by container
uids (the mail store is uid 5000) and find/du silently skip what they cannot
read. A copy run as hutao therefore loses mail without erroring. Root over
Tailscale SSH is what makes the copy complete — ssh root@<tailnet-ip> works
even though sudo on that box demands a password.
mail-state is 95 M and is the bulk of the mail data. It is not scratch: it holds
dovecot’s indexes and UIDVALIDITY, and rspamd’s trained bayes database. Skip it
and every IMAP client re-downloads everything and your spam filter starts from
zero.
Deliberately not transferred:
| Skipped | Why |
|---|---|
/srv/certbot/data | security.acme issues a fresh certificate; the certbot lineage layout is not the same and is not read any more |
/srv/dozzle/data | just users.yml, rendered from sops by dozzle-users.service |
/srv/caddy, caddy_caddy_* volumes | ACME state caddy no longer manages |
/srv/mailserver/data/dms/mail-logs | 11 M of logs |
roundcube/config | rendered from the Nix store |
postgres-data, serenity-discord-bot_postgres-data | both 0 bytes, 0 links — dead volumes |
deploy_* volumes, /srv/camofox-browser | expenses app and camofox, staying on the old box |
| docker build cache | 3.8 G of nothing |
~/.config/syncthing carries the node’s TLS keypair, which is its device ID.
Copy it and every paired device keeps working; regenerate it and you re-pair by
hand on every phone and laptop.
The Discord bot’s data
Not in a docker volume, which is why the two postgres volumes on the old box are
0 bytes: the bot talks to the host’s PostgreSQL 18.6, listening on
127.0.0.1:5432 and on the tailnet address. DATABASE_URL in the bot’s .env
points at database serenity_discord_bot, 8.4 MB, 9 tables — and only
user_stats has rows (258 of them). Everything else is empty schema.
Dumped read-only with pg_dump --no-owner --no-privileges, and held in two
places so it does not live only on the box being retired:
old box: ~/migration-dumps/serenity-<timestamp>.sql
workstation: ~/migration-dumps/serenity-<timestamp>.sql
The bot stayed on the old box through the cutover and was ported afterwards in
8bba9bb: host services.postgresql behind pgbouncer, the dump restored, and
its credentials in sops. --no-owner --no-privileges is what let the restore
land under a different role than the Ubuntu one that owned it.
The two Cloudflare tokens
There are two, with different jobs, and confusing them costs an hour:
| Where | Used by | Token ID |
|---|---|---|
secrets.yaml → cloudflare.api_token | lego / security.acme on the VPS | a13f8e29ba9ebe4201e5aef3d1723ec7 |
tofu/terraform.tfvars → cloudflare_api_token | tofu, from a workstation | c91db61e8b430d39e780fd5e6098c225 |
The VPS token carries an IP filter pinned to the old server, so DNS-01 fails
from anywhere else with 403 9109: Cannot use the access token from location: <ip>. security.acme then falls back to a self-signed certificate and starts
dependent services anyway, so caddy and DMS come up serving a placeholder and
nothing looks broken until you check the issuer:
ssh -p 2222 hutao@<host> sudo cat /var/lib/acme/hu-tao.dev/cert.pem \
| openssl x509 -noout -issuer
# issuer=CN=minica root ca … <- placeholder
# issuer=C=US, O=Let's Encrypt … <- real
Test the token directly rather than by triggering ACME — Let’s Encrypt caps failed validations at 5 per account per hostname per hour, and lego burns one per attempt:
curl -sS -H "Authorization: Bearer $TOKEN" \
'https://api.cloudflare.com/client/v4/zones?name=hu-tao.dev'
This resolves itself at cutover, when the box inherits the old server’s address.
Order of operations
Volumes must exist and be populated before their container first starts, otherwise docker creates them empty and the service initialises itself blank — forgejo would make a new instance, DMS an empty mail store. So: install, stop the containers, load the data, start.
Observed on the freshly installed box, and both are the CORRECT empty-state behaviour rather than faults to chase:
- forgejo serves HTTP 200 but
/api/v1/version404s andrepos: 0— it is a blank instance that has not been through its install wizard. Restoringforgejo_datais what makes it the real instance; do NOT click through the wizard first, or you create a second one. - DMS loops on
You need at least one mail account to start Dovecot (120s left...)and then exits, so systemd restarts it. Accounts live inpostfix-accounts.cfinside thedms_configvolume. Until that volume is restored there is no account, and DMS refuses to run Dovecot or Postfix — which is why ports 25/465/587/993 have no banner yet. Restoringdms_configresolves it; nothing needs fixing.
export SSH_AUTH_SOCK=/tmp/hutao-agent.sock # key is passphrase-protected
NEW=178.105.223.159
1. Install NixOS
nix run github:nix-community/nixos-anywhere -- \
--flake .#vps-hetzner \
--target-host root@$NEW \
--extra-files <staged age key dir>
vps-hetzner, not vps: it targets /dev/sda. The extra-files directory
carries /var/lib/sops-nix/key.txt at 0600 — without it the host boots unable
to decrypt anything, including its own root password.
2. Quiesce, then load data
ssh -p 2222 hutao@$NEW 'sudo systemctl stop "docker-*"'
Old box first, so nothing is written mid-copy:
ssh hutao@100.97.90.108 'sudo systemctl stop docker-services; \
cd /srv && for d in forgejo mailserver minecraft kuma navidrome; do \
(cd $d 2>/dev/null && sudo docker compose down); done'
Then, from the OLD box over the tailnet (traffic stays inside fsn1 rather than going via a workstation), for each row of the mapping:
sudo rsync -aHAX --numeric-ids --info=progress2 \
/srv/forgejo/data/ root@hu-tao:/var/lib/docker/volumes/forgejo_data/_data/
-aHAX --numeric-ids because uid/gid must survive verbatim: forgejo’s repos,
the mail store and grafana’s data are owned by container uids that mean nothing
in either host’s /etc/passwd.
Create each volume before writing into it:
ssh -p 2222 hutao@$NEW 'for v in forgejo_data dms_mail dms_state dms_config \
roundcube_db minecraft_data kuma_data navidrome_data grafana_data tempo_data; \
do sudo docker volume create $v; done'
3. Verify before cutover
DNS-01 works regardless of which IP the box holds, so certificates can be issued and the whole stack tested while the old box is still live.
ssh -p 2222 hutao@$NEW 'systemctl start docker-network-proxy; sudo systemctl start "docker-*"'
ssh -p 2222 hutao@$NEW 'systemctl is-system-running; systemctl --failed'
curl --resolve git.hu-tao.dev:443:$NEW https://git.hu-tao.dev/api/v1/version
curl --resolve status.hu-tao.dev:443:$NEW https://status.hu-tao.dev/api/entry-page
Check journalctl -u docker-mailserver shows the certificate loading from
/certs, and that forgejo lists the migrated repositories.
4. Cutover — the IP handover
Hetzner requires a server to be powered off to (un)assign a primary IP.
- Final delta rsync of the mapping rows (minutes, since only deltas move)
- Power off both servers
- Unassign
147045244(178.105.223.159) from163906050 - Unassign
134632948(167.233.24.58) from137766340 - Assign
134632948to163906050 - Assign
147045244to137766340— the old box stays online at the other address, so serenity-bot, camofox and the expenses app keep running - Power on both
smtp.hu-tao.dev follows 134632948 automatically; it is a property of the IP.
5. Rollback
Steps 2–7 in reverse. The old box is untouched, still holds its data, and has
delete and rebuild protection on. Rollback is ~5 minutes and costs nothing
but the swap.
(Historical: 137766340 has since been deleted, so this rollback no longer
exists.)
After it settled
All done. Kept as the checklist that was run:
- ✅ Attach firewall
11483636to163906050 - ✅ Enable
delete+rebuildprotection on163906050(tofu/server.tf) - ✅ Import into tofu: server, primary IPs, firewall, rDNS, and the Cloudflare
records (
tofu/imports.tf). Read the plan: abort if it showsdestroy and then createonhcloud_server - ✅ Empty
legacy_server_idsintofu/variables.tfonce the old box is retired - ✅ Port serenity-bot (
8bba9bb,modules/containers/serenity-bot.nix) - ✅ From then on deploys are
deploy .#vps, with auto-rollback on lockout
What changed in the port
Not a 1:1 translation. The deliberate departures:
-
Data lives in named docker volumes, never a bind-mounted host directory.
services.resticbacks up/var/lib/docker/volumeswholesale, so a service added later is covered the moment it declares a volume. A backup that has to be told about each new service is a backup that eventually stops covering one. The Ansible roles’ data guards (assertthatdata/worldexists before provisioning) have no equivalent and need none — there is no path to point at the wrong place. -
Configuration comes from the Nix store, read-only. Store files are 0444, which is what the tempo (uid 10001) and grafana (uid 472) permission failures in the Ansible setup were about; and a store path changes when its content does, so systemd recreates the container on a config-only change. That is what
recreate: alwayswas working around. -
Certificates are
security.acme, not a certbot container plus cron. The certificate is namedhu-tao.devwith the subdomains as SANs, so reordering the list cannot silently issue a second lineage the way certbot’s name-after-the-first--dbehaviour could. Note NixOS calls the keykey.pem, not certbot’sprivkey.pem, and DMS reads it withSSL_TYPE=manualrather than guessing from$SSL_DOMAIN. -
Images pin a release tag and no digest. The Ansible repo pinned
tag@sha256:…; the digests have since gone stale, and a release tag has been the more stable of the two in practice.:latestis still never used — for tempo it is a main-branch build that reports a version which was never released. -
Real credentials everywhere. Grafana’s anonymous-Admin access is off and the login comes from sops, as do dozzle’s bcrypt hash and minecraft’s RCON password. uptime-kuma is the exception, unavoidably: it has no environment variable or config file that seeds an admin account — the first visitor is prompted to create one and the route then closes. Create it immediately after the first deploy. searxng is the other exception, for the opposite reason: it has no concept of a user at all, so there is nothing to seed — caddy’s
basic_authis the entire access control and the hash lives in sops. Becausebasic_authruns a cost-14 bcrypt on every request, caddy’srate_limit(a compiled-in module) sits in front of it and returns 429 before the hash runs, so a password flood cannot become CPU exhaustion. It is in-process on purpose — see the fail2ban row in Failure modes for the forward-chain ban whose blast radius it avoids.Generate that hash with
mkpasswd, which is already on the host:mkpasswd -m bcrypt -R 14-R 14is not optional. mkpasswd defaults to cost 05 and caddy’s ownhash-passworduses 14, so the default silently produces a hash 512x cheaper to attack than the one caddy would have made. The$2b$prefix mkpasswd emits is fine — caddy verifies through golang.org/x/crypto/bcrypt, which records the minor version without validating it, so$2a$,$2b$and$2y$are interchangeable. -
Not ported:
camofox. It is a host systemd unit for an npm project checked out under/home, not a container service, and it depends on a tree this image does not create.serenity-botwas in the same position and was ported later as a container built from upstream’s Dockerfile; seemodules/containers/serenity-bot.nix.
Postmortems
Written when something broke badly enough that the fix is not obvious from the diff, and kept afterwards because the reasoning is the part that does not survive in git.
The house style: a timeline with real timestamps, impact measured rather than estimated, and a “what went badly” section that is allowed to be unflattering. A postmortem that only lists what went well is a press release.
| Date | Title | One-line cause |
|---|---|---|
| 2026-09-16 | A dependency window that took the box down for six minutes | a three-second transient in one non-critical container made deploy-rs abort, and the abort stopped every container |
The structural lesson from that one is in Failure modes: a failed activation leaves the box on the previous generation’s configuration, not its previous running state.
2026-09-16 — a dependency window that took the box down for six minutes
Impact: ~5m44s, 22:19:55–22:25:39 EEST. Every service except the two
minecraft worlds was down, mail included.
Trigger: a nix flake update deploy.
Root cause: a known, documented, deferred missing ordering dependency on
docker-forgejo-runner.service.
Detected by: kuma-check, correctly, within 13 seconds of its next tick.
1. What we set out to do
Clear all seven pending items on the Renovate Dependency Dashboard (#12) inside
one 30-minute maintenance window: three container bumps, one container major
(docker-mailserver 15.1.0 → 16.0.1), lockFileMaintenance, and two tofu
provider bumps.
Six of the seven landed. The seventh took the box down on its way in.
2. Timeline
All times EEST. The box logs UTC; subtract three hours.
| Time | Event |
|---|---|
| 22:00:40 | Cold-stop grafana + navidrome. Local tar, then restic a3534fdf. |
| 22:02:44 | Deploy 1 (gen 58): dozzle v11.1.0, grafana 13.2.2, navidrome 0.64.0. Clean, 40s. |
| 22:11:24 | Cold-stop mail. Local tar, then restic 63564755. |
| 22:12:43 | Deploy 2 (gen 59): docker-mailserver 16.0.1 + the opendkim gid fix. Clean, 37s. |
| 22:16:39 | Deploy 3 starts: flake.lock only. |
| 22:19:28 | New nixpkgs changes every unit → every container restarts at once. Runner stopped. |
| 22:19:52 | Runner starts, declares against https://git.hu-tao.dev/, caddy is not listening yet: fail to invoke Declare … connection refused. Exits 1. |
| 22:19:53 | deploy-rs samples unit state, sees one failed unit. |
| 22:19:55 | Deploy 3 aborts. De-activation stops every container. Outage begins. |
| 22:19:58 | The runner’s own Restart=always brings it up successfully — 3 seconds after the abort decision. |
| 22:20:08 | kuma-check fails: Failed to connect to status.hu-tao.dev:443. |
| 22:25:01 | kuma-check fails again. |
| 22:25:03 | systemctl restart docker (the documented recovery). |
| 22:25:09 | switch-to-configuration switch → Could not acquire lock. Fell back to starting units directly. |
| 22:25:39 | Mail answering. Outage ends. |
| 22:30:27 | kuma-check green again. |
| 22:33:16 | Deploy 4 (gen 60): flake.lock + the readiness gate. Clean, 40s. |
| 22:34:54 | Reboot complete on kernel 6.18.52. 38 seconds down. |
| ~22:38 | tofu apply: 0 added, 0 changed, 0 destroyed. |
3. Impact, measured
- Mail (25/465/587/993) refused connections for 5m44s. No bounces or deferrals are attributable to the window. That is expected but only weak evidence: while the container was down we logged nothing, because nothing reached us. Sending MTAs queue and retry on schedules measured in hours, and six minutes is far inside every one of them, so loss is very unlikely — but it is unprovable from our side.
- Everything behind caddy — forgejo, webmail, kuma, navidrome, searxng, pages — refused connections for the same period.
- Observability (grafana, tempo, dozzle) down, and kuma down, so the public status page could not report its own outage. This is the paradox Observability anticipates, and the mitigation worked: see §6.
- minecraft / minecraft2 restarted and came back on their own. The other fourteen containers did not.
4. Root cause
docker-forgejo-runner.service has a hard runtime dependency on caddy that is
nowhere expressed to systemd.
The runner’s first action on start is declaring itself against instanceUrl,
which is the public https://git.<domain>/ — deliberately, because the URL
is handed to job containers that sit on per-job networks where a proxy-network
container name does not resolve. That call therefore leaves the box and comes
back in through caddy. If caddy is not yet listening, the connect is refused and
the runner exits 1 rather than retrying in-process.
Because the runner declares no networks, modules/containers/default.nix
gives it no after/requires at all. Nothing has ever ordered it behind caddy.
This was known. Commit 9326a8c, 2026-09-13, closes with:
Separately, and NOT fixed here: the runner declares no
networks, so modules/containers/default.nix gives it no after/requires at all. Any future deploy that does restart docker.service reproduces this same failure. Worth an ordering dependency, as its own change.
That change was never written. This is the “future deploy”.
5. Was the cause “combining a system update with container updates”?
No — and the timeline is what rules it out. It is worth writing down because it was the first hypothesis, and it is a reasonable one.
The four updates were deployed in four separate deploys, not one. At the
moment of failure, deploy 3 contained flake.lock and the inert tofu lock file
and nothing else. The three container bumps were already live and untouched;
the mailserver major was already live and untouched. Removing them from the
branch entirely would have changed nothing about this failure.
What actually mattered is a property of the system update on its own: a new
nixpkgs changes the store path of every systemd unit, so switch-to-configuration
restarts every container simultaneously. That is the condition the runner
cannot survive, and it needs no container bump to arrive. The same thing
happened three days earlier for an unrelated reason — commit 523cbe1, enabling
the Actions cache, pinned the
docker daemon’s bip, which restarted docker.service, which stopped every
container — and produced the identical error string.
But the instinct behind the hypothesis is correct, and it is the real lesson.
The error was treating a system update as just another row on the same
checklist. The dashboard listed lock file maintenance directly beneath
update amir20/dozzle docker tag to v11.1.0, as though they were the same kind
of change. They are not, and the difference is blast radius:
| restarts | reboot | rollback | |
|---|---|---|---|
| a container tag bump | 1 unit | no | re-deploy, unless it migrated its own data |
lockFileMaintenance | every unit | yes, if the kernel moves | re-deploy, cleanly |
A window sized and sequenced for the first kind is not a window for the second. So: not “don’t combine them in a PR” — they combine in a PR fine, and did. It is “don’t deploy them in the same window, and never deploy the system one last.”
6. What went well
kuma-checkcaught it. The out-of-band timer failed at 22:20:08 and 22:25:01 and was green either side. The self-hosted-status-page paradox described in Observability was closed exactly as designed: kuma could not report its own outage, and the thing that noticed was the probe that lives outside it.- deploy-rs did its job on deploys 1, 2 and 4 — three clean activations with magic rollback confirmed, ~40s each.
- Backups were taken cold and verified before every migrating change. Two local tars and two tagged restic snapshots, all taken with the relevant containers stopped. None were needed. That is the correct outcome.
- The pre-flight on the mailserver major caught a real break before any downtime: opendkim’s gid moves 104 → 102 in v16, which would have made the box receive mail and silently refuse to send it.
- The documented recovery was correct.
systemctl restart dockerwas the right first move and it worked. See Failure modes.
7. What went badly
- A known landmine was left armed for three days.
9326a8cdiagnosed the exact failure, wrote down that it would recur, and deferred the fix. The deferral was reasonable in isolation and wrong in aggregate: the cost of the fix was a 59-line module addition, and the cost of not doing it was a six-minute full outage. - The system update was attempted in the last third of the window, after it had already been identified as needing its own. Thirteen minutes was not enough for a three-week nixpkgs move plus a reboot plus verification, and the window should have been re-scoped rather than squeezed.
- deploy-rs’s abort path leaves the box worse than either config. This is the most surprising finding here and deserves its own line: when activation fails, the de-activation stops the containers and the rollback does not restart them. The box does not return to the old generation’s running state — it returns to the old generation’s configuration, with nothing running. A 3-second transient in one non-critical unit therefore cost fifteen healthy services.
switch-to-configuration switchfailed withCould not acquire lockduring recovery, because deploy-rs still held it. Recovery had to fall back to starting units by hand. Worth knowing before the next incident.
8. What changed
-
b380c8afix(forgejo-runner): aforgejo-runner-ready.serviceoneshot, in the same shape asforgejo-runner-token, that probes/api/v1/versionand waits before the runner starts. Ordering alone would not have been enough — adocker-*unit counts as started when the container is created, not when the service inside it answers. It exits 0 on timeout deliberately: its job is to close the race, not to become a new way for a deploy to fail.It earned its place on the first cold boot after the reboot:
curl: (7) Failed to connect to git.hu-tao.dev:443 after 83 ms: Could not connect to server forgejo answered after 2 attempt(s)The first probe failed. Without the gate, that is the runner exiting 1.
-
The handbook’s Failure modes gains a row for the deploy-rs abort behaviour. (Done; it was
ARCHITECTURE.md§12 when this was written.)
9. Action items
| # | Action | Why | Status |
|---|---|---|---|
| 1 | Audit every container unit for an unexpressed runtime dependency on another container. The runner was found the hard way; it is unlikely to be the only one. | Same class of bug, same trigger. | Open. The runner was fixed in b380c8a; no other unit has been audited. |
| 2 | Deploy lockFileMaintenance alone, first in a window, never last. | §5. | Standing practice. |
| 3 | Decide whether a failed activation should de-activate at all. magicRollback protects against losing SSH; it is not obviously the right tool for one crashlooping non-critical unit, and its abort is more destructive than the failure it responds to. | §7. | Open. magicRollback and autoRollback are unchanged in flake.nix. |
| 4 | Confirm the healthchecks.io grace period is short enough that two missed pings actually page. The probe failed correctly; whether that produced an alert is configured outside this repo. | §6 is only half-verified. | Unverified; configured outside this repo. |
| 5 | Create postmaster@hu-tao.dev. | RFC 5321 §4.5.1. | Done in 0944f97, with abuse@ alongside. |