Hardening from two incidents on 2026-06-06/07: Schema drift (morning): drizzle-kit push exits 0 even when it aborts on an interactive prompt it can't render in CI, so cd-staging's set -euo pipefail never fired and a month of staging schema drift accumulated silently until new code hit missing columns. All three drizzle push call sites now tee output and fail the deploy on any 'Error:' line. Production outage (overnight, ~9h): cd-infra and cd-apps had no failure handling at all. A network-option change stopped postgres for a network recreation that then deadlocked on containers from other compose projects holding trails-shared; the script carried on, the stack stayed down. Both scripts now run set -euo pipefail and gate on container health at the end (postgres+journal for cd-infra, journal+planner for cd-apps) so a deploy that leaves the stack down is a red X, not a shrug. docs/deployment.md gains the cross-project manual procedure for network-changing deploys — the CD workflows only manage their own compose project and cannot apply those safely. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
6 KiB
Deployment runbook
trails.cool runs on two Hetzner hosts. This document covers what an
operator needs to know beyond the CLAUDE.md summary.
Hosts
| Role | Host | IPs | SSH |
|---|---|---|---|
| Flagship (Cloud) | trails.cool |
public + 10.0.0.2 (vSwitch) |
ssh -i ~/.ssh/trails-cool-deploy root@trails.cool |
| BRouter (Dedicated) | ullrich.is |
public 176.9.150.227 + 10.0.1.10 (vSwitch) |
ssh -i ~/.ssh/trails-brouter-deploy -p 2232 trails@ullrich.is |
Both hosts are in fsn1 (Falkenstein) and joined on Hetzner vSwitch
#80672 (VLAN 4000). The flagship's Terraform (infrastructure/terraform/)
owns the Cloud Network + subnets + server attachment. The dedicated
host's VLAN sub-interface is configured out-of-band via netplan
(/etc/netplan/60-trails-vswitch.yaml), because the Robot side isn't
in the Hetzner Cloud API.
BRouter host — first-time provisioning
See infrastructure/brouter-host/README.md. The short version:
# As root (one-time firewall allowances):
ufw allow in on enp4s0.4000 from 10.0.0.2 to any port 17777 proto tcp \
comment 'trails brouter via flagship vSwitch'
ufw allow in on enp4s0.4000 from 10.0.0.2 to any port 8080 proto tcp \
comment 'trails brouter cadvisor via flagship vSwitch'
# As the trails user:
cd ~/brouter # created by the first cd-brouter deploy
./download-segments.sh # ~10 GB, a few minutes on a good connection
docker compose pull
docker compose up -d
Secrets rotation
Tokens (including BROUTER_AUTH_TOKEN):
- Generate:
openssl rand -base64 32 - Edit:
SOPS_AGE_KEY_FILE=~/.config/sops/age/keys.txt sops infrastructure/secrets.app.env - Commit + push + merge →
cd-appsredeploys the Planner with the new token. - Touch anything under
infrastructure/brouter-host/(or rungh workflow run cd-brouter.yml) →cd-brouterredeploys the Caddy sidecar with the new token. - Brief overlap window where Planner sends new token while Caddy still checks the old value. Both redeploys should complete within a minute of each other; in the worst case a few Planner requests get 403 and retry.
SOPS on macOS looks for the age key at ~/Library/Application Support/sops/age/keys.txt by default. If yours lives under XDG-standard ~/.config/sops/age/keys.txt, set SOPS_AGE_KEY_FILE as above or export it in your shell rc.
Cutover procedure (flagship BRouter → dedicated host)
This is how BROUTER_URL gets flipped. Do it once the dedicated host
is provisioned, segments are seeded, and the compose project is up.
- Pre-flight:
curl -sfH "X-BRouter-Auth: $(sops -d infrastructure/secrets.app.env | grep ^BROUTER_AUTH_TOKEN= | cut -d= -f2-)" http://10.0.1.10:17777/brouter?lonlats=11.58,48.13\|11.59,48.14\&profile=trekking\&alternativeidx=0\&format=gpxfrom the flagship. Expect 200 with GPX. Then curl without the header — expect 403. - Wire the token without flipping the URL. Edit SOPS:
sops infrastructure/secrets.app.env— theBROUTER_AUTH_TOKENis already in there. IfBROUTER_URLisn't in SOPS, skip; the compose has a default. Merge. Planner redeploys; it now sends the header to the flagship BRouter (which ignores it). - Flip the URL. In SOPS, add
BROUTER_URL=http://10.0.1.10:17777. Merge.cd-appsredeploys the Planner. - Monitor. Grafana "BRouter (dedicated host)" dashboard +
brouter_request_duration_secondson the Overview board. Watch for 30 minutes. - Rollback (if needed): remove the
BROUTER_URLline from SOPS (falls back to the flagship default). Merge; redeploy. The flagship container is still warm during the soak window. - Decommission flagship BRouter (after 48 h of clean metrics): remove the
brouter:service +./segmentsvolume frominfrastructure/docker-compose.yml. Merge.cd-infrarestarts without BRouter. Reclaim ~2 GB of segment volume on the flagship.
Full restart (flagship)
gh workflow run cd-infra.yml -f restart_all=true
Restarts every flagship service. Does NOT touch the BRouter host.
Network-changing deploys (flagship)
Changing options on an existing Docker network (enable_ipv6, subnets,
drivers) requires Docker to recreate the network — and a network can
only be recreated when no container from any compose project is
attached. The flagship has three+ projects sharing trails-shared
(production, persistent staging, every PR preview), and the CD workflows
only manage their own project, so a network-option change shipped through
cd-infra alone WILL deadlock mid-deploy and can leave production down
(this is exactly the 2026-06-06/07 outage: postgres stopped for the
recreation, the deploy failed on the held network, nothing restarted it
for ~9 hours).
The manual procedure, in order, on the flagship:
cd /opt/trails-cool
# 1. Free trails-shared: down every preview + persistent staging
docker compose ls --filter name=trails-pr- --format json # enumerate previews
docker compose -f docker-compose.staging.yml -p trails-pr-<N> --env-file staging-pr-<N>.env down
docker compose -f docker-compose.staging.yml -p trails-staging --env-file staging.env --profile persistent down
# 2. Recreate networks via the production project
docker compose --env-file app.env down
docker compose --env-file app.env up -d
# 3. Bring staging + previews back
docker compose -f docker-compose.staging.yml -p trails-staging --env-file staging.env --profile persistent up -d
docker compose -f docker-compose.staging.yml -p trails-pr-<N> --env-file staging-pr-<N>.env up -d
# 4. Verify
docker network inspect trails-cool_default trails-shared --format '{{.Name}} ipv6={{.EnableIPv6}}'
curl -sf https://trails.cool/api/health && curl -sf https://staging.trails.cool/api/health
Plan it as a short maintenance window (~2–3 min downtime); don't ship network-option changes expecting the workflows to apply them.
cd-brouter manual trigger
gh workflow run cd-brouter.yml
Useful to redeploy the BRouter host after token rotation or a config
change, without needing a real source change under
infrastructure/brouter-host/.