The 2026-06-07 disk-full outage (flagship 100%, postgres crash-looping on its pidfile, prod + staging down): cd-apps DOES prune images after deploys, but it had been failing early all day (migration bug), so its prune never ran — while ~10 staging/preview deploys kept pulling fresh images with no prune of their own. - cd-staging: prune superseded layers (until=1h guard against racing in-flight pulls) after persistent staging and preview deploys. - disk-maintenance.yml: NEW daily scheduled prune (04:30 UTC) that also FAILS when the disk is still ≥85% after pruning — a redundant alert channel for exactly the case where the Grafana disk alert drowns in other noise, as it did during the incident. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| cd-apps.yml | ||
| cd-brouter.yml | ||
| cd-infra.yml | ||
| cd-staging.yml | ||
| ci.yml | ||
| dependabot-dedupe.yml | ||
| disk-maintenance.yml | ||
| staging-cleanup.yml | ||
| update-visual-snapshots.yml | ||