The journal/planner deploy in cd-apps.yml does `docker compose up -d journal planner`, which stops the old container and starts the new one — Caddy keeps forwarding requests during the ~10–30s gap and returns 502s. The caddy-502-rate alert (threshold > 0 for 2m) correctly trips, every time. Two production changes plus a long-broken workflow detail: - infrastructure/Caddyfile — add `lb_try_duration 30s` / `lb_try_interval 250ms` to the journal and planner reverse_proxy blocks. Caddy now holds and retries the upstream for up to 30s during a restart instead of 502'ing immediately. Real outages (upstream unreachable longer than 30s) still 502 and the alert still fires for those. - infrastructure/grafana/provisioning/alerting/alerts.yml — add a comment documenting why caddy-502-rate stays at threshold > 0: with lb_try_duration in front of it, the alert no longer conflates "deploy in flight" with "real outage." - .github/workflows/cd-apps.yml — fix a long-silent bug: the Grafana deploy-annotation step was reading GRAFANA_SERVICE_TOKEN from `.env`, but the secrets file we scp to /opt/trails-cool is named `app.env`. The token check failed silently and the curl was being skipped on every deploy. Switching to `app.env` so deploys actually annotate Grafana. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| cd-apps.yml | ||
| cd-brouter.yml | ||
| cd-infra.yml | ||
| ci.yml | ||
| dependabot-dedupe.yml | ||