veza/docs/runbooks at ccf3e64d9a1246119e2c3144faafcf7d80811eb1 - senke/veza

senke/veza

History

senke 7d92820a9c docs(runbooks): expand INCIDENT_RESPONSE + GRACEFUL_DEGRADATION stubs Both files were ~15-25 lines of bullet points — fine as a placeholder, useless under stress at 03:00 when the on-call has never seen Veza misbehave before. Expanded both to the same depth as db-failover.md / redis-down.md / rabbitmq-down.md so the on-call has an actual runbook to follow. INCIDENT_RESPONSE.md (15 → 208 lines) * "First 5 minutes" triage : ack → annotation → 3 dashboards → failure-class matrix → declare-if-stuck. Aligns with what an on-call actually does when paged. * Severity ladder (SEV-1/2/3) with response-time and communication norms — replaces the implicit "everything is SEV-1" the bullet points suggested. * "Capture evidence before mitigating" block with the four exact commands (docker logs, pg_stat_activity, redis bigkeys, RMQ queues) the postmortem will want. * Mitigation patterns per failure class (API down, DB down, storage failure, webhook failure, DDoS, performance), each pointing at the deep-dive runbook for the specific recipe. * "After mitigation" : status page, comm pattern, postmortem schedule by severity, runbook update policy. * Tools section with the bookmark-able URLs (Grafana, Tempo, Sentry, status page, HAProxy stats, pg_auto_failover monitor, RabbitMQ console, MinIO console). GRACEFUL_DEGRADATION.md (25 → 261 lines) * Quick-lookup matrix of every backing service × user-visible impact × severity × deep-dive runbook. Lets the on-call read one row instead of paging through six docs. * Per-service section detailing what still works and what fails : Postgres primary/replica, Redis master/Sentinel, RabbitMQ, MinIO/S3, Hyperswitch, Stream server, ClamAV, Coturn, Elasticsearch (called out as the v1.0 orphan it is). * `/api/v1/health/deep` documented as the canary surface, with a sample response shape so operators know what `degraded` looks like before they see it. * "Adding a new degradation mode" section with the 4-step recipe (this file, /health/deep, alert annotation, FAIL-SOFT/FAIL-LOUD code comment) so future maintainers keep the docs in sync as the surface evolves. These two files now match the depth of the alert-specific runbooks ; no more "open the runbook, find 15 lines, panic" path. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>		2026-05-01 04:13:55 +02:00
..
game-days	docs(release): game day #2 prod session + v2.0.0-rc1 release notes (W6 Day 28)	2026-04-29 15:44:32 +02:00
api-availability-slo-burn.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
api-latency-slo-burn.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
cert-expiring-soon.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
db-failover.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
DEPLOYMENT.md	chore(release): v0.961 — Playbook (runbooks déploiement, rollback, incident)	2026-03-02 19:09:46 +01:00
disk-full.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
GRACEFUL_DEGRADATION.md	docs(runbooks): expand INCIDENT_RESPONSE + GRACEFUL_DEGRADATION stubs	2026-05-01 04:13:55 +02:00
INCIDENT_RESPONSE.md	docs(runbooks): expand INCIDENT_RESPONSE + GRACEFUL_DEGRADATION stubs	2026-05-01 04:13:55 +02:00
payment-success-slo-burn.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
rabbitmq-down.md	fix(scripts,docs): game-day prod safety guards + rabbitmq-down runbook	2026-04-30 22:32:05 +02:00
redis-down.md	feat(observability): SLO burn-rate alerts + 7 runbook stubs (W2 Day 10)	2026-04-28 01:30:34 +02:00
ROLLBACK.md	chore(release): v0.961 — Playbook (runbooks déploiement, rollback, incident)	2026-03-02 19:09:46 +01:00
SECRET_ROTATION.md	chore(release): v0.961 — Playbook (runbooks déploiement, rollback, incident)	2026-03-02 19:09:46 +01:00