ci(deploy): fail loudly — set -e, drizzle output guards, health gates
Hardening from two incidents on 2026-06-06/07: Schema drift (morning): drizzle-kit push exits 0 even when it aborts on an interactive prompt it can't render in CI, so cd-staging's set -euo pipefail never fired and a month of staging schema drift accumulated silently until new code hit missing columns. All three drizzle push call sites now tee output and fail the deploy on any 'Error:' line. Production outage (overnight, ~9h): cd-infra and cd-apps had no failure handling at all. A network-option change stopped postgres for a network recreation that then deadlocked on containers from other compose projects holding trails-shared; the script carried on, the stack stayed down. Both scripts now run set -euo pipefail and gate on container health at the end (postgres+journal for cd-infra, journal+planner for cd-apps) so a deploy that leaves the stack down is a red X, not a shrug. docs/deployment.md gains the cross-project manual procedure for network-changing deploys — the CD workflows only manage their own compose project and cannot apply those safely. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
007845cb59
commit
f790da2ed3
4 changed files with 99 additions and 3 deletions
21
.github/workflows/cd-infra.yml
vendored
21
.github/workflows/cd-infra.yml
vendored
|
|
@ -64,6 +64,11 @@ jobs:
|
|||
username: root
|
||||
key: ${{ secrets.DEPLOY_SSH_KEY }}
|
||||
script: |
|
||||
# Abort on first failure. The 2026-06-06/07 outage: a network
|
||||
# recreation stopped postgres, a later step failed, and the
|
||||
# deploy left production down for ~9h while the job's partial
|
||||
# progress looked plausible. Fail fast, verify health at the end.
|
||||
set -euo pipefail
|
||||
cd /opt/trails-cool
|
||||
|
||||
# .env was placed by the SCP step (decrypted app + infra secrets)
|
||||
|
|
@ -93,6 +98,22 @@ jobs:
|
|||
|
||||
docker compose ps
|
||||
|
||||
# Gate on the stack actually being up: postgres healthy and —
|
||||
# since an infra restart bounces the app containers' database —
|
||||
# journal back to healthy too. A deploy that leaves either down
|
||||
# must fail loudly (see the 2026-06-06/07 outage).
|
||||
for ctr in trails-cool-postgres-1 trails-cool-journal-1; do
|
||||
for i in $(seq 1 36); do
|
||||
status=$(docker inspect -f '{{.State.Health.Status}}' "$ctr" 2>/dev/null || echo missing)
|
||||
[ "$status" = "healthy" ] && break
|
||||
sleep 5
|
||||
done
|
||||
if [ "$status" != "healthy" ]; then
|
||||
echo "$ctr did not become healthy (last status: $status)"
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
# Annotate deploy in Grafana
|
||||
GRAFANA_TOKEN=$(grep GRAFANA_SERVICE_TOKEN .env | cut -d= -f2-)
|
||||
if [ -n "$GRAFANA_TOKEN" ]; then
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue