Stop the caddy-502-rate alert firing on every deploy

The journal/planner deploy in cd-apps.yml does `docker compose up -d
journal planner`, which stops the old container and starts the new
one — Caddy keeps forwarding requests during the ~10–30s gap and
returns 502s. The caddy-502-rate alert (threshold > 0 for 2m)
correctly trips, every time.

Two production changes plus a long-broken workflow detail:

- infrastructure/Caddyfile — add `lb_try_duration 30s` /
  `lb_try_interval 250ms` to the journal and planner reverse_proxy
  blocks. Caddy now holds and retries the upstream for up to 30s
  during a restart instead of 502'ing immediately. Real outages
  (upstream unreachable longer than 30s) still 502 and the alert
  still fires for those.
- infrastructure/grafana/provisioning/alerting/alerts.yml — add a
  comment documenting why caddy-502-rate stays at threshold > 0:
  with lb_try_duration in front of it, the alert no longer
  conflates "deploy in flight" with "real outage."
- .github/workflows/cd-apps.yml — fix a long-silent bug: the
  Grafana deploy-annotation step was reading
  GRAFANA_SERVICE_TOKEN from `.env`, but the secrets file we scp
  to /opt/trails-cool is named `app.env`. The token check failed
  silently and the curl was being skipped on every deploy.
  Switching to `app.env` so deploys actually annotate Grafana.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Ullrich Schäfer 2026-04-26 11:51:13 +02:00
parent 153f133093
commit 5c4b6fd9af
3 changed files with 30 additions and 4 deletions

View file

@ -28,7 +28,17 @@
output stdout
format json
}
reverse_proxy journal:3000
reverse_proxy journal:3000 {
# During an `apps` deploy the journal container is briefly down
# (~1030s) while compose swaps containers. Without these,
# Caddy returns 502 immediately and the `caddy-502-rate` alert
# trips on every deploy. With them, Caddy holds and retries
# against the upstream for up to 30s — restart becomes
# invisible to clients. A real outage longer than 30s still
# 502s and correctly trips the alert.
lb_try_duration 30s
lb_try_interval 250ms
}
}
www.{$DOMAIN:trails.cool} {
@ -53,5 +63,9 @@ planner.{$DOMAIN:trails.cool} {
output stdout
format json
}
reverse_proxy planner:3001
reverse_proxy planner:3001 {
# Same rationale as the journal block — see the comment there.
lb_try_duration 30s
lb_try_interval 250ms
}
}

View file

@ -206,6 +206,12 @@ groups:
annotations:
summary: "BRouter host metrics scrape has been failing for 2+ minutes — the dedicated host, vSwitch, or cAdvisor may be down"
# The threshold here is intentionally `> 0` for 2m — *any*
# sustained 502 stream is real. Deploy-time restarts no longer
# produce 502s thanks to `lb_try_duration` on Caddy's reverse
# proxy (see `infrastructure/Caddyfile`); if 502s appear here
# it means the upstream has been unreachable for longer than
# Caddy's retry window, which is a genuine outage.
- uid: caddy-502-rate
title: Caddy 502 errors detected
condition: B