Task group 3 of federation-hardening. There was no blocklist of any kind;
the only lever against a hostile instance was an IP/host block in Caddy.
- New `federation_blocked_instances` table (domain PK, reason,
created_at). Additive → created by drizzle-kit push.
- `federation-blocklist.server.ts`: exact-host matching —
`isBlockedDomain`, `isBlockedIri` (unparseable IRI ⇒ treated as
blocked), and `filterBlockedDomains` for batch recipient filtering.
- Enforced at all three boundaries (spec: federation-operations
"Instance blocklist"):
- inbox — each of the 4 listeners silently drops a blocked actor's
activity (202, no error oracle) before dedup/side effects;
- delivery enqueue — `enqueueActivityDeliveries` filters blocked
recipients in one batch query;
- outbox poll / actor fetch — `pollRemoteActor` refuses a blocked host
up front (`skipped: "blocked instance"`), before any network.
- Operator procedure (SQL insert/list/delete) documented in the
deployment runbook's federation section.
- Tests: unit (hostOfIri) + integration against real Postgres covering
the helper and the delivery + outbox boundaries; inbox uses the same
tested isBlockedIri primitive.
Note: the inbox-drop *counter* (federation_inbox_dropped_total{reason})
lands with the other metrics in task 4.2; this commit is the enforcement.
Verified: db + journal typecheck + lint clean; drizzle-kit push creates
the table; blocklist integration tests green against real Postgres;
journal unit suite 357 passing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.8 KiB
Deployment runbook
trails.cool runs on two Hetzner hosts. This document covers what an
operator needs to know beyond the CLAUDE.md summary.
Hosts
| Role | Host | IPs | SSH |
|---|---|---|---|
| Flagship (Cloud) | trails.cool |
public + 10.0.0.2 (vSwitch) |
ssh -i ~/.ssh/trails-cool-deploy root@trails.cool |
| BRouter (Dedicated) | ullrich.is |
public 176.9.150.227 + 10.0.1.10 (vSwitch) |
ssh -i ~/.ssh/trails-brouter-deploy -p 2232 trails@ullrich.is |
Both hosts are in fsn1 (Falkenstein) and joined on Hetzner vSwitch
#80672 (VLAN 4000). The flagship's Terraform (infrastructure/terraform/)
owns the Cloud Network + subnets + server attachment. The dedicated
host's VLAN sub-interface is configured out-of-band via netplan
(/etc/netplan/60-trails-vswitch.yaml), because the Robot side isn't
in the Hetzner Cloud API.
BRouter host — first-time provisioning
See infrastructure/brouter-host/README.md. The short version:
# As root (one-time firewall allowances):
ufw allow in on enp4s0.4000 from 10.0.0.2 to any port 17777 proto tcp \
comment 'trails brouter via flagship vSwitch'
ufw allow in on enp4s0.4000 from 10.0.0.2 to any port 8080 proto tcp \
comment 'trails brouter cadvisor via flagship vSwitch'
# As the trails user:
cd ~/brouter # created by the first cd-brouter deploy
./download-segments.sh # ~10 GB, a few minutes on a good connection
docker compose pull
docker compose up -d
Secrets rotation
Tokens (including BROUTER_AUTH_TOKEN):
- Generate:
openssl rand -base64 32 - Edit:
SOPS_AGE_KEY_FILE=~/.config/sops/age/keys.txt sops infrastructure/secrets.app.env - Commit + push + merge →
cd-appsredeploys the Planner with the new token. - Touch anything under
infrastructure/brouter-host/(or rungh workflow run cd-brouter.yml) →cd-brouterredeploys the Caddy sidecar with the new token. - Brief overlap window where Planner sends new token while Caddy still checks the old value. Both redeploys should complete within a minute of each other; in the worst case a few Planner requests get 403 and retry.
SOPS on macOS looks for the age key at ~/Library/Application Support/sops/age/keys.txt by default. If yours lives under XDG-standard ~/.config/sops/age/keys.txt, set SOPS_AGE_KEY_FILE as above or export it in your shell rc.
Cutover procedure (flagship BRouter → dedicated host)
This is how BROUTER_URL gets flipped. Do it once the dedicated host
is provisioned, segments are seeded, and the compose project is up.
- Pre-flight:
curl -sfH "X-BRouter-Auth: $(sops -d infrastructure/secrets.app.env | grep ^BROUTER_AUTH_TOKEN= | cut -d= -f2-)" http://10.0.1.10:17777/brouter?lonlats=11.58,48.13\|11.59,48.14\&profile=trekking\&alternativeidx=0\&format=gpxfrom the flagship. Expect 200 with GPX. Then curl without the header — expect 403. - Wire the token without flipping the URL. Edit SOPS:
sops infrastructure/secrets.app.env— theBROUTER_AUTH_TOKENis already in there. IfBROUTER_URLisn't in SOPS, skip; the compose has a default. Merge. Planner redeploys; it now sends the header to the flagship BRouter (which ignores it). - Flip the URL. In SOPS, add
BROUTER_URL=http://10.0.1.10:17777. Merge.cd-appsredeploys the Planner. - Monitor. Grafana "BRouter (dedicated host)" dashboard +
brouter_request_duration_secondson the Overview board. Watch for 30 minutes. - Rollback (if needed): remove the
BROUTER_URLline from SOPS (falls back to the flagship default). Merge; redeploy. The flagship container is still warm during the soak window. - Decommission flagship BRouter (after 48 h of clean metrics): remove the
brouter:service +./segmentsvolume frominfrastructure/docker-compose.yml. Merge.cd-infrarestarts without BRouter. Reclaim ~2 GB of segment volume on the flagship.
Full restart (flagship)
gh workflow run cd-infra.yml -f restart_all=true
Restarts every flagship service. Does NOT touch the BRouter host.
Network-changing deploys (flagship)
Changing options on an existing Docker network (enable_ipv6, subnets,
drivers) requires Docker to recreate the network — and a network can
only be recreated when no container from any compose project is
attached. The flagship has three+ projects sharing trails-shared
(production, persistent staging, every PR preview), and the CD workflows
only manage their own project, so a network-option change shipped through
cd-infra alone WILL deadlock mid-deploy and can leave production down
(this is exactly the 2026-06-06/07 outage: postgres stopped for the
recreation, the deploy failed on the held network, nothing restarted it
for ~9 hours).
The manual procedure, in order, on the flagship:
cd /opt/trails-cool
# 1. Free trails-shared: down every preview + persistent staging
docker compose ls --filter name=trails-pr- --format json # enumerate previews
docker compose -f docker-compose.staging.yml -p trails-pr-<N> --env-file staging-pr-<N>.env down
docker compose -f docker-compose.staging.yml -p trails-staging --env-file staging.env --profile persistent down
# 2. Recreate networks via the production project
docker compose --env-file app.env down
docker compose --env-file app.env up -d
# 3. Bring staging + previews back
docker compose -f docker-compose.staging.yml -p trails-staging --env-file staging.env --profile persistent up -d
docker compose -f docker-compose.staging.yml -p trails-pr-<N> --env-file staging-pr-<N>.env up -d
# 4. Verify
docker network inspect trails-cool_default trails-shared --format '{{.Name}} ipv6={{.EnableIPv6}}'
curl -sf https://trails.cool/api/health && curl -sf https://staging.trails.cool/api/health
Plan it as a short maintenance window (~2–3 min downtime); don't ship network-option changes expecting the workflows to apply them.
Federation runbook
Federation (ActivityPub via Fedify) is gated by FEDERATION_ENABLED=true
per environment. Currently: staging on, production off. The full
change history and design live in openspec/changes/social-federation/.
Enabling on an environment
- Ensure
FEDERATION_KEY_ENCRYPTION_KEYis insecrets.app.env(SOPS) — keypair generation fails closed in production without it. - Set
FEDERATION_ENABLED=truein the environment's compose env (seecd-staging.ymlfor the staging pattern). - First boot with the flag on enqueues
backfill-user-keypairs(generates RSA keys for existing users — idempotent) and registers the federation jobs (deliver-activity,poll-remote-actor,poll-remote-outboxes,federation-kv-sweep). - Smoke:
curl -s "https://<domain>/.well-known/webfinger?resource=acct:<user>@<domain>"(200 for a public user, 404 otherwise) and/nodeinfo/2.1.
FEDERATION_LOG_LEVEL=debug turns on Fedify's per-request signature
diagnostics (staging runs debug during soaks; dial back to info).
Key rotation
Rotating FEDERATION_KEY_ENCRYPTION_KEY requires re-encrypting every
users.private_key_encrypted (decrypt with old, encrypt with new) —
there is no script yet; write one against
apps/journal/app/lib/federation-keys.server.ts primitives when first
needed. Rotating a single user's keypair: NULL their key columns and
let the backfill regenerate (remotes pick the new key up from the actor
document; deliveries signed with the old key fail until then).
Abuse monitoring
- Inbox is rate-limited 60 req/5 min per source instance
(
federationSourceHost); 429s show in journal logs. - Outbound is paced (1 req/s delivery, 1 req/5 s polling per host).
- Watch
docker logs <journal>forfedify·federationlines; Loki picks them up.
Blocking an instance
Blocking a domain makes it inert in both directions: its inbound
activities are silently dropped (a 202, no error oracle), we send it no
deliveries, and we never fetch its actors/outboxes. Matching is
exact-host. v1 management is a SQL insert/delete against
journal.federation_blocked_instances (the table is the seam a future
admin UI can sit on); protocol/moderation semantics are in
FEDERATION.md.
-- Block a hostile instance (exact host, no scheme, no path):
INSERT INTO journal.federation_blocked_instances (domain, reason)
VALUES ('bad.example', 'spam / harassment')
ON CONFLICT (domain) DO NOTHING;
-- List current blocks:
SELECT domain, reason, created_at FROM journal.federation_blocked_instances ORDER BY created_at DESC;
-- Unblock:
DELETE FROM journal.federation_blocked_instances WHERE domain = 'bad.example';
The block takes effect immediately (checked per-request/per-job, no cache). Already-queued deliveries to the domain drain from Fedify's queue; new fan-outs filter it out. An IP/host block in Caddy remains the harder emergency lever for a flood that shouldn't reach the app at all.
Troubleshooting deliveries
Hard-won checklist from the 2026-06-06/07 soak (details in
openspec/changes/social-federation/design.md):
- Test both IP families from inside the journal container before
suspecting signatures (
family: 6vs4— dangling A records and v6-only instances are common; our containers have IPv6 since #466/#468). - Remote shows the post count but no posts → remotes never backfill outbox history; only pushed or individually-fetched objects appear.
- Delivered but invisible, no errors anywhere → check the remote's
tombstonestable; a previously-retracted URI is refused forever. - Mastodon gives deliveries a 10s read timeout — anything synchronous and slow in the inbox path breaks federation silently.
- Actor changes (profile fields, keys) need the remote to re-fetch:
tootctl accounts refresh user@domainon a Mastodon you control.
cd-brouter manual trigger
gh workflow run cd-brouter.yml
Useful to redeploy the BRouter host after token rotation or a config
change, without needing a real source change under
infrastructure/brouter-host/.