Stop the caddy-502-rate alert firing on every deploy
The journal/planner deploy in cd-apps.yml does `docker compose up -d journal planner`, which stops the old container and starts the new one — Caddy keeps forwarding requests during the ~10–30s gap and returns 502s. The caddy-502-rate alert (threshold > 0 for 2m) correctly trips, every time. Two production changes plus a long-broken workflow detail: - infrastructure/Caddyfile — add `lb_try_duration 30s` / `lb_try_interval 250ms` to the journal and planner reverse_proxy blocks. Caddy now holds and retries the upstream for up to 30s during a restart instead of 502'ing immediately. Real outages (upstream unreachable longer than 30s) still 502 and the alert still fires for those. - infrastructure/grafana/provisioning/alerting/alerts.yml — add a comment documenting why caddy-502-rate stays at threshold > 0: with lb_try_duration in front of it, the alert no longer conflates "deploy in flight" with "real outage." - .github/workflows/cd-apps.yml — fix a long-silent bug: the Grafana deploy-annotation step was reading GRAFANA_SERVICE_TOKEN from `.env`, but the secrets file we scp to /opt/trails-cool is named `app.env`. The token check failed silently and the curl was being skipped on every deploy. Switching to `app.env` so deploys actually annotate Grafana. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
153f133093
commit
5c4b6fd9af
3 changed files with 30 additions and 4 deletions
|
|
@ -28,7 +28,17 @@
|
|||
output stdout
|
||||
format json
|
||||
}
|
||||
reverse_proxy journal:3000
|
||||
reverse_proxy journal:3000 {
|
||||
# During an `apps` deploy the journal container is briefly down
|
||||
# (~10–30s) while compose swaps containers. Without these,
|
||||
# Caddy returns 502 immediately and the `caddy-502-rate` alert
|
||||
# trips on every deploy. With them, Caddy holds and retries
|
||||
# against the upstream for up to 30s — restart becomes
|
||||
# invisible to clients. A real outage longer than 30s still
|
||||
# 502s and correctly trips the alert.
|
||||
lb_try_duration 30s
|
||||
lb_try_interval 250ms
|
||||
}
|
||||
}
|
||||
|
||||
www.{$DOMAIN:trails.cool} {
|
||||
|
|
@ -53,5 +63,9 @@ planner.{$DOMAIN:trails.cool} {
|
|||
output stdout
|
||||
format json
|
||||
}
|
||||
reverse_proxy planner:3001
|
||||
reverse_proxy planner:3001 {
|
||||
# Same rationale as the journal block — see the comment there.
|
||||
lb_try_duration 30s
|
||||
lb_try_interval 250ms
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -206,6 +206,12 @@ groups:
|
|||
annotations:
|
||||
summary: "BRouter host metrics scrape has been failing for 2+ minutes — the dedicated host, vSwitch, or cAdvisor may be down"
|
||||
|
||||
# The threshold here is intentionally `> 0` for 2m — *any*
|
||||
# sustained 502 stream is real. Deploy-time restarts no longer
|
||||
# produce 502s thanks to `lb_try_duration` on Caddy's reverse
|
||||
# proxy (see `infrastructure/Caddyfile`); if 502s appear here
|
||||
# it means the upstream has been unreachable for longer than
|
||||
# Caddy's retry window, which is a genuine outage.
|
||||
- uid: caddy-502-rate
|
||||
title: Caddy 502 errors detected
|
||||
condition: B
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue