Companion to PR #358 which moved the staging compose ports off 3100/3101
(Loki conflict on the vSwitch). The Caddyfile staging blocks have the
upstream ports baked in so they need bumping too — already applied
manually on the flagship to unblock staging; this lands it in-repo so it
survives the next cd-infra deploy.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Two bugs in the staging-environments rollout:
1. Staging containers published on 127.0.0.1 are unreachable from the
production Caddy container, which connects via the Docker bridge IP
(host.docker.internal:host-gateway resolves to the bridge, not
loopback). Bind to 0.0.0.0 instead — Hetzner Cloud firewall blocks
ports 3000+ from the public internet, so it stays internal-only.
2. The trailing 'docker compose ps' diagnostic in deploy-staging /
deploy-preview was missing --env-file staging.env, so compose
failed env interpolation and the job exited non-zero even when the
actual deploy succeeded.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Implements the staging-environments OpenSpec change. Persistent staging at
staging.trails.cool / planner.staging.trails.cool deploys from main; PR
opens get a journal-only preview at pr-<N>.staging.trails.cool that shares
the persistent planner. cd-staging.yml builds tagged images, manages
per-PR Postgres databases and Caddyfile snippets, evicts the oldest
preview at the cap of 3, and tears everything down on PR close.
staging-cleanup.yml runs weekly to sweep orphaned previews.
DNS records (staging + *.staging A/AAAA) already applied to production
via tofu.
Caddy approach: per-PR Caddyfile snippets imported from /etc/caddy/sites/
and reloaded on each PR event — no wildcard / on-demand TLS, no router
service. Production compose gains a trails-shared network for the staging
project to reach Postgres, and host.docker.internal on Caddy so it can
reverse-proxy staging containers published on the host loopback.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The journal/planner deploy in cd-apps.yml does `docker compose up -d
journal planner`, which stops the old container and starts the new
one — Caddy keeps forwarding requests during the ~10–30s gap and
returns 502s. The caddy-502-rate alert (threshold > 0 for 2m)
correctly trips, every time.
Two production changes plus a long-broken workflow detail:
- infrastructure/Caddyfile — add `lb_try_duration 30s` /
`lb_try_interval 250ms` to the journal and planner reverse_proxy
blocks. Caddy now holds and retries the upstream for up to 30s
during a restart instead of 502'ing immediately. Real outages
(upstream unreachable longer than 30s) still 502 and the alert
still fires for those.
- infrastructure/grafana/provisioning/alerting/alerts.yml — add a
comment documenting why caddy-502-rate stays at threshold > 0:
with lb_try_duration in front of it, the alert no longer
conflates "deploy in flight" with "real outage."
- .github/workflows/cd-apps.yml — fix a long-silent bug: the
Grafana deploy-annotation step was reading
GRAFANA_SERVICE_TOKEN from `.env`, but the secrets file we scp
to /opt/trails-cool is named `app.env`. The token check failed
silently and the curl was being skipped on every deploy.
Switching to `app.env` so deploys actually annotate Grafana.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The old home was an h1, subtitle, and two auth buttons — visitors
arriving at trails.cool had no idea what it was, and self-hosted
instances had nothing interesting to land on.
The new home is one layout that serves both audiences:
- Hero: a product-describing h1 ("Federated outdoor journal") +
tagline. The site name already lives in the top banner and nav brand,
so the h1 carries the pitch instead of triplicating "trails.cool".
- Auth CTAs: Register (blue) + Sign In (outlined) get primary weight.
Planner demotes to a small "Or try the Planner without an account →"
link below them — it's a nice escape hatch but not the goal.
- Marketing cards: on the flagship only, a 2x2 grid of emoji-icon cards
for Planner / Journal / Federation / Ownership, matching the Planner
home's visual weight. `IS_FLAGSHIP=true` toggles it. Self-hosters get
a "Powered by trails.cool — about the project" link back to the
flagship instead, so they don't have to write their own marketing.
- Public feed: reverse-chronological list of the 20 most recent public
activities on the instance, with owner + distance + date. Empty state
for fresh installs.
Plumbing:
- `listRecentPublicActivities(limit)` in activities.server.ts joins
users so the feed can render "by <displayName>" in one query.
- `IS_FLAGSHIP` env wired through docker-compose.yml; cd-apps and
cd-infra both write `IS_FLAGSHIP=true` alongside `DOMAIN=trails.cool`,
so self-hosters (who deploy without these workflows) default to off.
Specs:
- `activity-feed`: new "Instance-wide public activity feed" requirement
so the visibility contract is pinned (public in feed, unlisted +
private are not).
- `journal-landing`: new small spec capturing the hero / marketing /
feed layout contract and the `IS_FLAGSHIP` toggle, so future
instance-branding or empty-state changes have a stable anchor.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The notification policy and contact points live in
`infrastructure/grafana/provisioning/alerting/alerts.yml`, so
UI-created contact points can't be attached to provisioned alerts.
Provision the Pushover integration instead.
Rather than add a second contact point and do the fan-out in the
notification tree (Grafana's tree only fires one receiver per match
without duplication workarounds), put both email + pushover under a
single `default` contact point. Every alert now fans out to both.
Secrets go to SOPS `secrets.infra.env`. The Grafana container reads
`PUSHOVER_API_TOKEN` and `PUSHOVER_USER_KEY_ULLRICH` from env; the
provisioning YAML references them via `$VAR` and Grafana substitutes
at startup, so no secret lands in git.
Priority 1 on firing (bypass quiet hours), 0 on resolve. User can
tweak in the YAML later.
Post-deploy: delete the UI-created "Pushover Ullrich" contact point
in Grafana to avoid a duplicate. It's unused (no policy references it)
but shows up in the contact-point list.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Users land via GitHub OAuth gated by the trails-cool org allowlist, so
Viewer-only is more restrictive than we actually want — no Explore
access, no ad-hoc Loki/Prometheus queries. Promote the default role to
Editor so new sign-ins get it automatically.
Note: GF_USERS_AUTO_ASSIGN_ORG_ROLE only applies to new users on first
sign-in. Existing Viewers (including the single current user) need a
one-off promotion via the Grafana admin UI.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After PR #293 flipped BROUTER_URL to the dedicated host and Grafana +
manual smoke tests confirmed clean routing (US route worked, request
rate/latency/memory/logs all green), there's no reason to keep the
in-tree flagship brouter container warm any longer.
- Remove the `brouter:` service and its `./segments` bind mount from
`infrastructure/docker-compose.yml`.
- Remove `depends_on: brouter` from the planner service.
- Tighten the BROUTER_URL and BROUTER_AUTH_TOKEN env wiring from
`${…:-default}` to `${…:?message}` so a missing SOPS value fails the
compose up loudly instead of silently pointing at the removed service
(or missing auth).
- Tick tasks 5.5, 7.5, 8.4, 9.3 in the OpenSpec change; 9.4 (archive)
is the last remaining step once this merges.
Post-merge operator step: `docker image prune -f` on the flagship to
reclaim the brouter image.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When a user rapidly moves waypoints in the Planner, BRouter's
thread-priority watchdog kills older in-flight requests to free a thread
for newer ones, surfacing to the user as "operation killed by
thread-priority-watchdog" errors.
Two fixes:
1. Bump the BRouter thread pool on the dedicated host from 4 to 16.
`BROUTER_THREADS` is now a Dockerfile ENV (default 4 for
flagship/CI/dev) overridden to 16 in `infrastructure/brouter-host/`.
The host has 32 GB of RAM and idle cores; 4 threads was a flagship-era
constraint that no longer applies.
2. Cancel the previous in-flight `/api/route` request when a new one
starts. `use-routing.ts` now keeps an `AbortController` in a ref,
aborts it at the top of each `computeRoute`, and bails quietly on
`AbortError` so the newer call drives state.
Together these eliminate both the "real" contention (at the BRouter
level) and the wasted work (BRouter computing a route the client no
longer cares about).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds BROUTER_URL=http://10.0.1.10:17777 (vSwitch) to SOPS so the Planner
starts routing through the dedicated host's Caddy sidecar instead of the
in-tree flagship BRouter. Flagship BRouter stays warm for the 48h
soak/rollback window.
Pre-flight from the flagship confirmed the new host answers 200 with the
auth header and 403 without.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three fixes after the first cd-brouter run on the dedicated host:
1. **BRouter 1.7.8 → 1.7.9** (`docker/brouter/Dockerfile`). Planet RD5
segments on brouter.de are now version 11; 1.7.8's `lookups.dat`
is v10, causing `lookup version mismatch (old rd5?)
lookups.dat=10 E10_N45.rd5=11` on every route request.
2. **cd-brouter.yml: docker login to ghcr.io before pull**.
ghcr.io/trails-cool/brouter is private, and the dedicated host's
Docker daemon isn't logged in by default. Extract DEPLOY_GHCR_TOKEN
from SOPS at runner side, pass to the SSH step via envs, and
`docker login` before `docker compose pull`. Credential is
`::add-mask::`-ed so it doesn't show in logs.
3. **Drop custom healthcheck** on the brouter service. The image
strips wget/curl post-build, and /bin/sh in the base doesn't
support /dev/tcp, so there's no in-image way to do an HTTP probe.
Real health is observed via Caddy's upstream 502 behavior on
outage and the Planner-side `brouter_request_duration_seconds`
metric. caddy's `depends_on` drops from service_healthy to
service_started.
End-to-end verified on the dedicated host after applying the compose
fix manually:
- Caddy enforces auth: 403 without header, proxies with.
- BRouter 1.7.9 will resolve the segment-version error once the image
is rebuilt.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Lands sections 6 and 8 of relocate-brouter-to-dedicated-host on top
of the host compose in #291.
## Observability (section 6)
- **Prometheus**: new `brouter-cadvisor` job scraping
`10.0.1.10:8080` over the vSwitch, labeled `host="brouter"`.
- **cAdvisor sidecar** on the dedicated host's compose
(`--docker_only --whitelisted_container_labels=trails.cool.service`)
so metrics only cover trails containers, never the operator's
unrelated workloads on the shared box.
- **Promtail sidecar** on the dedicated host, Docker SD with relabel-
drop on missing `trails.cool.service` label, pushing to flagship
Loki at `http://10.0.0.2:3100/loki/api/v1/push`.
- **Flagship compose**: Loki now publishes port 3100 on the vSwitch
IP only (10.0.0.2:3100) — Hetzner Cloud firewall blocks it from
the public internet.
- **Grafana dashboard**: `brouter.json` — scrape up/down, request
rate (from Planner-side `brouter_request_duration_seconds`),
p50/p95/p99, container memory/CPU, Loki logs panel.
- **Alert**: `brouter-scrape-down` fires on
`up{job="brouter-cadvisor"} < 1 for 2m`; `noDataState: Alerting`
so a total scrape failure still pages.
Operator needs one UFW rule on the dedicated host for the cAdvisor
port — documented in `infrastructure/brouter-host/README.md`.
## Documentation (section 8)
- `CLAUDE.md` — hosts table + updated deployment table with SSH
targets per workflow; BRouter host SSH is `-p 2232 trails@...`,
different key.
- `docs/architecture.md` — Hosting section rewritten to cover both
hosts, vSwitch boundary, and the observability-scoping rationale
for the shared dedicated host.
- `docs/deployment.md` (new) — full operator runbook: host layout,
first-time BRouter provisioning, SOPS rotation (including the
macOS `SOPS_AGE_KEY_FILE` gotcha), cutover procedure with
rollback, manual workflow triggers.
Task 8.4 (infrastructure/README.md) skipped: that file doesn't
exist and the ground is covered by brouter-host/README.md +
docs/deployment.md.
## Validation
- `docker compose config` on both the flagship and brouter-host
compose files — both validate.
- `pnpm typecheck`, `pnpm lint`, `pnpm test` — all clean (full
turbo cache hits; no code changes in this commit).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Lands sections 3-5 of the relocate-brouter-to-dedicated-host change:
everything needed to run BRouter on the dedicated Hetzner Robot host
and have the Planner talk to it with the shared-secret header. Does
NOT flip the cutover — the flagship BRouter stays warm during soak.
## BRouter host compose (section 3)
New `infrastructure/brouter-host/` — a standalone compose project that
runs as the `trails` user on `ullrich.is`:
- `docker-compose.yml` — brouter + caddy sidecar. BRouter has no host
port; caddy binds only to `10.0.1.10:17777` (vSwitch IP). Every
service explicitly overrides the host's default Loki logging driver
to `json-file` so logs don't leak to the operator's personal Loki.
- `Caddyfile` — single-purpose reverse proxy that requires
`X-BRouter-Auth: ${BROUTER_AUTH_TOKEN}` on every request. `auto_https
off` (vSwitch-only); default access log format omits request
headers, so the token is never written to disk.
- `download-segments.sh` — crawls brouter.de, pulls planet-wide RD5
tiles via `wget -N` (incremental). Idempotent, safe to cron.
- `README.md` — one-shot provisioning + token rotation + rollback
notes.
`docker/brouter/Dockerfile` is patched to honor `JAVA_OPTS` (was
hardcoded `-Xmx1024M` in CMD). Default keeps the flagship's current
heap; compose on the dedicated host overrides to `-Xmx8g` for planet
scale on a 32 GB box.
## Planner shared-secret header (section 4)
`apps/planner/app/lib/brouter.ts`:
- Module-level guard: throws at startup in production if
`BROUTER_AUTH_TOKEN` is unset.
- `authHeaders()` helper (reads env at call time, so tests can
`vi.stubEnv` without module reset).
- Header attached on both `computeRoute` (per-segment) and
`computeSegmentGpx`.
3 new unit tests cover header attachment + the no-token path.
`infrastructure/docker-compose.yml` passes `BROUTER_AUTH_TOKEN` to
the Planner service, and makes `BROUTER_URL` overridable via SOPS so
the cutover is a one-variable flip.
## cd-brouter workflow (section 5)
Rewritten to deploy to the dedicated host:
- SSH as `trails@${BROUTER_DEPLOY_HOST}` on port
`${BROUTER_DEPLOY_SSH_PORT}` (2232) using
`${BROUTER_DEPLOY_SSH_KEY}`.
- Decrypts SOPS, extracts ONLY `BROUTER_AUTH_TOKEN` into a `.env`
file, scp'd alongside the compose project.
- `paths:` trigger now includes `infrastructure/brouter-host/**`.
- Segment download is NOT run here — first-time seed is a manual
operator step (multi-hour). Routine re-runs are cron-able on the
dedicated host.
- Grafana annotation step preserved (reaches flagship Grafana as
before).
## What's NOT here
- `brouter:` service on the flagship is intentionally left in place
(removed in section 7.5 after the 48 h soak window post-cutover).
- Observability (section 6) — Prometheus scrape + Loki shipping from
the dedicated host — comes in a follow-up PR.
- Cutover itself (section 7) — flip `BROUTER_URL`, verify, remove the
flagship brouter — is an operator action gated on first-time
provisioning + smoke testing.
## Verification
`pnpm typecheck && pnpm lint && pnpm test` all clean; planner build
passes (the CI regression from #286 was fixed in #290).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the shared-secret header value consumed by the Planner
(as X-BRouter-Auth outbound) and by the BRouter host's Caddy
sidecar (enforced inbound). Lives in secrets.app.env only;
cd-brouter.yml will also read from this file, making it the
single source of truth. secrets.infra.env is intentionally NOT
updated — cd-infra no longer touches BRouter after the move.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an OpenSpec change scoping the relocation of BRouter from the
co-located flagship (cx23, 4 GB RAM, 40 GB SSD, Europe segments only) to
a dedicated Hetzner Robot host in the same datacenter, with private
connectivity over Hetzner vSwitch #80672 (VLAN 4000).
This first PR only lays the network prerequisite:
- Terraform: a Hetzner Cloud Network (10.0.0.0/16) with a cloud subnet
(10.0.0.0/24) hosting the flagship at 10.0.0.2, and a vSwitch subnet
(10.0.1.0/24) bridged to Robot VLAN 4000. The dedicated host's VLAN
sub-interface (10.0.1.10 on enp4s0.4000) is configured out-of-band via
netplan and is not Terraform-managed.
- lifecycle { ignore_changes = [user_data] } on the flagship server to
prevent the Hetzner provider's post-1.45 user_data hash drift from
triggering a spurious full-server replacement on unrelated applies.
- OpenSpec change with proposal, design, specs (delta for
brouter-integration / infrastructure / observability /
security-hardening), and tasks; Section 1 (pre-flight) is checked off
with operator notes.
Verification: ping both directions across the vSwitch is 0% loss,
sub-ms latency; dedicated host's VLAN config persists across reboot
(verified ~60 s to restore private reachability).
Follow-up PRs will land the BRouter host compose project, Planner
shared-secret header, CD workflow retarget, observability wiring, and
the cutover.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Public Overpass instances go bad without warning — we just lived
through a day where overpass.private.coffee (our single upstream)
returned 504s after 60s while lz4.overpass-api.de served the same
query in 0.6s. Survive that by trying a list of upstreams in order
until one responds usefully.
Server changes:
- OVERPASS_URLS env — comma-separated list, tried in order. Falls
back to OVERPASS_URL for backward compat, then to a built-in
default pair (lz4.overpass-api.de, overpass-api.de).
- Per-upstream timeout of 10s via AbortSignal.timeout, so a saturated
instance fails over in seconds, not minutes.
- Client abort propagation via AbortSignal.any — if the browser
cancels (e.g. user panned the map), the in-flight upstream fetch is
aborted too. Stops wasting upstream capacity on requests nobody
wants.
- Rate-limited-body detection (Overpass returns 200 with a
`rate_limited` marker when throttling) triggers failover.
- Per-upstream labels on overpass_upstream_requests_total and
overpass_upstream_duration_seconds so the dashboard can break out
health + latency by instance. Histogram buckets extended to 30s.
Compose: OVERPASS_URLS passthrough with the default pair hard-wired,
overridable via SOPS.
Tests: 8 cases covering the new fetchWithFailover helper — happy
path, 5xx failover, rate_limited failover, network error failover,
timeout tagging, all-upstreams-fail, client abort stops the loop,
per-attempt latency observation.
Follow-ups left out of scope:
- Client-side debounce (UI concern).
- Moving `.observe()` after body read for accuracy in measurement.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
pg-boss v11 renamed every concatenated-lowercase timestamp column
(`createdon`, `completedon`) to snake_case (`created_on`,
`completed_on`). Our service-health dashboard and failed-job alert
still referenced the v10 names, so after the v12 upgrade those
queries error out silently.
Fix the three affected queries:
- alerts.yml: failed-job alert
- service-health.json: "Recent Failed Jobs" table
- service-health.json: "Job Queue — Completed/hour" time series
The `state` / `name` / `output` columns stay the same, and the `state`
enum still accepts `'failed'` / `'completed'` literals.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Nine panels at trails-demo-bot covering:
- Current synthetic route/activity counts + time-since-last-walk
- Gauge time series over the configured retention window
- Walks per day + distance distribution
- Loki log stream filtered to `demo-bot` messages
- Recent walks table with name, distance, ascent
Provisioned via the existing grafana/dashboards directory — no extra
wiring needed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sets DEMO_BOT_ENABLED=true plus DEMO_BOT_RETENTION_DAYS and
DEMO_BOT_REGION in the SOPS env file so the journal worker starts
the generate + prune jobs and bootstraps the Bruno demo user on
next deploy.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds DEMO_BOT_PERSONA env var (inline JSON or file:<path>) so
self-hosted instances can replace Bruno-in-Berlin with their own
demo account identity and content pools. No-op for trails.cool — the
built-in Bruno persona remains the default.
- DemoPersona type + Zod schema validation
- loadPersona() cached at boot; falls back to default on any failure
- ensureDemoUser throws DemoPersonaUsernameClashError when the
persona username is already a real user; server declines to
schedule demo jobs for that process
- isDemoUser flag moved from hardcoded "bruno" check to loader-supplied
boolean computed from the running persona
- docs/demo-persona.md explains the schema + operator flow
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds a background demo user (`bruno`) plus two pg-boss jobs that generate
public trekking routes and activities in inner Berlin and prune them
after 14 days. Disabled by default everywhere; flip `DEMO_BOT_ENABLED`
only in prod.
- Schema: `synthetic` boolean on routes + activities
- `ensureDemoUser()` idempotent insert + badge on /users/bruno
- Generation gate: 07-21 Berlin local, p=0.12, 40-route cap in 14d
- Initial 4-walk backfill on first enable; retention via daily prune
- Prometheus gauges `demo_bot_synthetic_routes_total` / `_activities_total`
- Unit + DB-gated integration + E2E tests
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The original formulas used clamp_min(denominator, 1) to avoid divide-
by-zero, but that inflates the denominator when request rate is < 1/s.
Real example from prod: ~0.04 req/s, all successful, displayed as
1.02% health (red) instead of ~100%.
Switch to (numerator or vector(0)) / denominator:
- Missing "hit" series (no cache hits yet) → numerator resolves to 0
instead of empty set, so the stat shows 0% instead of "No data".
- Zero traffic (empty denominator) → whole expr is empty, Grafana
renders "No data", which is the correct behaviour for a health
stat when nothing's happening.
- Non-zero low traffic → ratio is true ratio, not artificially
depressed by the clamp.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds four metrics surfaced via the existing /metrics endpoint:
- overpass_cache_events_total{result=hit|miss|coalesced}
- overpass_cache_size (gauge)
- overpass_upstream_duration_seconds (histogram)
- overpass_upstream_requests_total{status} — covers 2xx/4xx/5xx and
an explicit "error" label for network-level failures
Grafana dashboards/planner.json gets a new row:
- Upstream health (5m 2xx success ratio, red <90%, green >99%)
- Cache hit ratio
- Cache size
- Upstream p95 latency
Plus timeseries panels for cache event breakdown, upstream status
over time, and p50/p95/p99 upstream latency.
The success-rate formula uses clamp_min(..., 1) on the denominator so
it doesn't NaN when traffic is zero.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds apps/planner/app/routes/api.overpass.ts — a same-origin, rate-limited
server-side proxy that forwards Overpass QL to the upstream endpoint
configured via OVERPASS_URL (default: overpass.private.coffee).
Motivation:
- private.coffee's best-practices require a meaningful User-Agent
identifying the project. Browsers cannot set User-Agent on fetch()
(forbidden header), so the request has to originate server-side.
- Collaborative sessions commonly have N clients pan/zoom the same map,
so the same bbox query arrives multiple times within seconds.
- Setting up the proxy now lets us swap OVERPASS_URL to a self-hosted
Overpass later without client changes.
What the proxy does:
- Sets User-Agent "trails.cool Planner (https://trails.cool; legal@trails.cool)"
- Enforces same-origin via the Origin header
- Rate-limits per client IP (120 req/min)
- In-memory LRU cache of upstream responses keyed on the form-encoded body
(TTL 10 min, max 200 entries)
- Coalesces concurrent misses for the same key so N simultaneous clients
in one session incur exactly one upstream call
Client change: apps/planner/app/lib/overpass.ts POSTs to /api/overpass
instead of iterating over public Overpass endpoints. Fallback list
removed; resilience is now the proxy's responsibility.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add @trails-cool/jobs package wrapping pg-boss for durable background job
execution using the existing PostgreSQL database. Wire up planner session
expiry as the first scheduled job (hourly, 7-day TTL) — this was previously
dead code that never ran. Add placeholder worker to journal for future
Komoot import and federation jobs.
Infrastructure: grant grafana_reader access to pgboss schema, add job queue
health panels to Service Health dashboard, add failed job alert.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fix blind spot where Caddy 502 errors were invisible: the dashboards and
alerts queried caddy_http_response_duration_seconds (upstream responses only),
missing 502s that Caddy generates itself. Switch to
caddy_http_request_duration_seconds (server-level, all responses).
Add Journal Grafana dashboard with: 502 rate, response codes, request rate
by route, latency percentiles, container restarts/memory/CPU, Node.js event
loop lag and heap, and Loki log panels for errors and Caddy 5xx entries.
Add color coding to Caddy status code panel (green=2xx, blue=3xx, yellow=4xx,
red=5xx). Add log-based error rate panel to the overview dashboard.
New alerts: container restart loop, PostgreSQL connections > 80, application
crash log detection (Loki), and Caddy 502 rate.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Switch journal Docker health check from `node -e "fetch(...)"` to `curl -sf`,
matching the planner. The Node.js process spawn was heavy and prone to timeout
under memory pressure, causing false health check failures and 502s.
Upgrade Promtail from static file scraping to Docker service discovery so logs
get `service` and `container` labels. Parse Pino JSON logs to extract the
`level` label, enabling Loki queries filtered by log level.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Dashboard: switch error rate panel from app-level http_request_duration
(never observed) to Caddy's caddy_http_response_duration_seconds_count.
Journal: add custom server.ts (matching planner pattern) that observes
http_request_duration_seconds on every request. Replaces react-router-serve
with a custom Node HTTP server that handles /api/health, /api/metrics,
static assets, and request timing. Removes redundant React Router routes
for health and metrics.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Caddy Response Status Codes: caddy_http_responses_total doesn't exist,
use caddy_http_response_duration_seconds_count which has the code label
- pg_stat_statements: enable extension in init script, add custom queries
config for postgres-exporter, remove invalid datname filter from query
- Add postgres/queries.yml to cd-infra SCP sources
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Loki was running but had no log shipper — no labels or logs were
visible in Grafana. Adds Promtail to scrape all Docker container
logs and push them to Loki.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
WAHOO_CLIENT_ID, WAHOO_CLIENT_SECRET, and WAHOO_WEBHOOK_TOKEN were in
SOPS secrets but never passed through to the journal service in
docker-compose.yml, causing the OAuth authorize URL to have an empty
client_id.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Without this, GitHub OAuth creates a new Grafana user on every login
instead of linking to the existing account by email.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add SSH hardening to user_data for future server rebuilds:
- PasswordAuthentication no (key-only access)
- X11Forwarding no (headless server)
- AllowTcpForwarding no (no SSH tunnels needed)
- AllowAgentForwarding no (no agent forwarding needed)
- MaxAuthTries 3 (reduced from 6)
- ClientAliveInterval 300 + ClientAliveCountMax 2 (clean up dead sessions)
Already applied manually to the existing server via
/etc/ssh/sshd_config.d/99-hardening.conf.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The alert was firing with DatasourceNoData because zero 5xx responses
means the numerator returns empty (not 0). Fixed by:
- Adding `or vector(0)` so zero 5xx returns 0% instead of empty
- Setting `noDataState: OK` as a safety net
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Grafana file-based provisioning expects flat dashboard JSON, not the
API format with a "dashboard" wrapper. The 4 wrapped dashboards were
silently failing to load — only business.json (already flat) worked.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The type:dashboard field was causing Grafana to filter annotations
by dashboard scope. Our deploy annotations are global (no dashboardUID),
so the target.type:tags is the only filter needed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The annotation config was missing the target block that tells
Grafana to filter by tags. Without it, the annotation query
was silently ignored.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The error rate formula (errors/total*100) produces NaN when there's
zero traffic (0/0). Grafana treats NaN as a firing condition.
Use clamp_min on the denominator to return 0% instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The Planner's save-to-journal callback makes a cross-origin request to
trails.cool from planner.trails.cool. Add the Journal's origin to the
Planner's connect-src CSP directive.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The threshold expressions were missing the 'expression' field referencing
the query refId, causing 'no variable specified for refId C' errors.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Panels:
- Stat cards: total users, routes, activities, active planner sessions
- Time series: cumulative user growth, route growth
- Bar charts: signups/day, routes/day, planner sessions/day
- Table: top 20 routes by distance
All queries use the grafana_reader PostgreSQL datasource.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Create grafana_reader role with SELECT-only access on all schemas
- Init script runs on postgres first boot (docker-entrypoint-initdb.d)
- Grafana PostgreSQL datasource provisioned with read-only credentials
- Enables SQL queries in dashboards (user count, routes, etc.)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>