# Incident: Production Database Wiped by `migrate:fresh`

**Date:** 2026-08-07
**Host:** `petav3-dev` (serves `https://wk.propertylab.com.my`)
**Database:** `petav3wk`, on a **remote** MySQL host
**Severity:** Total data loss on the primary database, partially recovered
**Cause:** An assistant (Claude Code) ran `php artisan migrate:fresh` believing environment variables had redirected it to a scratch database. They had not.

---

## Summary

While building the VR360 feature, a migration needed verifying against an empty database — a check GUIDELINES §7 explicitly requires. The command run was:

```bash
DB_DATABASE=petav3_migcheck DB_HOST=127.0.0.1 DB_USERNAME=root \
  php artisan migrate:fresh --force
```

Every one of those `DB_*` variables was **silently ignored**. Laravel read its connection from `bootstrap/cache/config.php`, which had been regenerated minutes earlier by the same session. `migrate:fresh` therefore dropped every table on the live remote database and rebuilt the schema empty: 449 migrations, zero rows.

The command was written to be safe, ran without a permission prompt, and reported success. Nothing in the output indicated it had hit production — the log line reads `Dropping all tables ... 3s DONE` with no host named.

---

## Root cause

**A cached Laravel config outranks environment variables.**

`php artisan config:cache` compiles every config file into `bootstrap/cache/config.php`. Once that file exists, `config('database.connections.mysql.*')` is read from it, and `env()` is never consulted. Prefixing a command with `DB_DATABASE=…` sets a variable that nothing reads.

The session created this exact hazard for itself. A new config file (`config/vr_bake.php`) was invisible until `config:cache` was run, so it was run — which disarmed the redirection mechanism used ten minutes later. The two actions were never connected.

The decisive detail: `migrate:fresh` **drops tables first, then migrates**. When its target is wrong, there is no partial failure and no recoverable state — the destruction is the first thing it does.

---

## Why the existing warning did not prevent it

`CLAUDE.local.md` says, unambiguously:

> Hostname is `petav3-dev`, but `.env` has `APP_ENV=production` and it serves the real domain. **Treat it as live: deploys cause real downtime.**

That file was loaded and read at the start of the session. A memory note recording *this precise failure mode* — a deployed config cache making CLI commands hit the live database — was also in context.

**Both were advisory. Neither was enforcing.**

An instruction in a context window has to be recalled and applied by a fallible reader at the exact moment of action — competing for attention against a twelve-task build with momentum behind it. A permission rule, a framework guard or a shell hook is evaluated by a machine on every invocation, indifferent to what anyone believes at the time.

When an advisory instruction and a mechanical permission disagree, **the mechanical one wins every time.** Here the allowlist said "yes, run it," and it won.

This is the central lesson. The remediation below therefore contains no new documentation. Docs did not fail for lack of clarity; they failed because they were the wrong category of control.

---

## Contributing factors

### 1. A wildcard permission pre-approved the command

`.claude/settings.local.json` contained:

```json
"Bash(php artisan *)"
```

That glob covers `migrate:fresh`, `migrate:refresh`, `migrate:reset` and `db:wipe`. There was no confirmation prompt because every artisan command had already been granted, permanently.

### 2. Laravel ships a guard for exactly this, and it was off

`Illuminate\Support\Facades\DB::prohibitDestructiveCommands()` exists in the installed framework and is registered nowhere in the application. `.env` already sets `APP_ENV=production`, so a single line in `AppServiceProvider::boot()` would have thrown instead of dropping.

### 3. Every backup was empty — and had been for weeks

`scripts/deploy-update.sh` reads `DB_DATABASE`, `DB_USERNAME` and `DB_PASSWORD` out of `.env` before calling `mysqldump` — **and never reads `DB_HOST`.** The dump therefore connects to localhost, where the database does not exist, and writes an empty stream.

> **Correction (2026-08-08).** An earlier revision of this document said `mysqldump` "exits cleanly, so the deploy log reports a successful backup". That is wrong, and the truth is worse. The log carries `WARN DB backup failed — continuing` at *every* occurrence — the failure was detected, reported, and then **ignored by design**. The deploy went on to migrate a database it had just failed to back up, and the warning scrolled past in a long log nobody re-reads. A silent failure is a bug; a failure you announce and then proceed past is a decision.
>
> Fixing this also uncovered a **second, independent** fault. With `-h` supplied the dump reached the server and still died after 367 bytes: `Couldn't execute 'FLUSH TABLES': Access denied … RELOAD`. In MySQL 8, `--single-transaction` requires the **global** `RELOAD` privilege, and this account holds `ALL PRIVILEGES` on its own schema but only `USAGE ON *.*`. So even a correctly-addressed backup would have produced nothing. Both are fixed in `95178e45`; a real dump is 188 MB.

```
-rw-rw-r-- 20 Jul 26 15:20 db-20260726-152040.sql.gz
-rw-rw-r-- 20 Jul 27 01:26 db-20260727-012658.sql.gz
   … 56 files, every one 20 bytes — the size of an empty gzip stream …
-rw-rw-r-- 20 Aug  7 09:42 db-20260807-094247.sql.gz   ← hours before the incident
```

**This is what turned a mistake into a catastrophe.** A backup was taken the same morning. It contained nothing. The only usable dump was a manual one from 26 July.

---

## Impact

| | |
|---|---|
| Tables dropped | All (449 migrations re-applied to an empty schema) |
| Rows lost at the moment of the incident | Every row in the primary database |
| Site behaviour | Served an empty schema; continued accepting writes |
| Recovered to | **2026-07-28 18:03** for catalogue tables; current for `users` / `leads` |
| Permanently lost | Catalogue changes made between 28 July and 7 August that were not reproducible from disk |

The restore is **mixed-vintage**, which matters when reasoning about it:

- `catalog_projects`, `catalog_project_sources`, `catalog_floor_plans` — newest row **28 July**
- `users`, `leads` — current
- Schema and the `migrations` table — from the accidental run, not the restore (it lists a migration authored the same day, which cannot have existed on 28 July)
- `permissions` — **79 rows against 82 constants in code**, because `backfill_missing_permissions` re-ran against an empty database and had nothing to backfill

---

## Recovery

Two things made recovery possible, neither of them the backup system.

**1. The scraped payloads were on disk, not in the database.** The wipe hit MySQL; the filesystem was untouched.

```
storage/app/catalogue/inbox/edgeprop-nl/processed/   8 files → 2,748 distinct records
storage/app/catalogue/inbox/iproperty/processed/     2 files →   151 distinct records
storage/app/catalogue/scrape/edgeprop-nl/            state.json + full HTTP cache
```

Every record was a complete normalised payload (`fields`, `floor_plans`, `media`, `raw`) — not a stub needing a second fetch. The iProperty crawl alone had cost **4,980 ScraperAPI credits**; re-scraping would have spent them again for data already on disk.

**2. PropertySifu's source was never on this database.** It reads from the `reference` connection — a different database on a different host — which still held all 301 rows.

Recovery was therefore **re-ingestion, not re-scraping**: no external calls, no credits.

| Feed | Result |
|---|---|
| PropertySifu | 36 created + 265 updated = **301** ✅ |
| iProperty | **151 created**, 0 failed ✅ |
| EdgeProp New Launches | re-ingesting 8 artifacts, oldest first (target ~2,735) |

### A latent bug surfaced during recovery

Two records (`53009` SJCC East One, `52014` Residensi Puncak Mertajam) are dropped by every ingestion run — **neither created, nor updated, nor failed.** `catalog_sync_runs.stats` has buckets for `created`, `updated` and `failed` and **none for "skipped"**, so a run that loses records reports itself as a clean success.

Worth fixing independently of this incident: a silent drop is indistinguishable from a complete run.

---

## Recommendations

Ordered by value. None of these are documentation changes.

### A. Register Laravel's destructive-command guard — 2 lines

```php
// app/Providers/AppServiceProvider.php — boot()
DB::prohibitDestructiveCommands($this->app->isProduction());
```

Blocks `migrate:fresh`, `migrate:refresh`, `migrate:reset` and `db:wipe` whenever `APP_ENV=production` — which this box already sets. Works regardless of who runs the command: assistant, human, script, or a future teammate who clones the repo and doesn't know the history.

**This alone would have prevented the incident.**

### B. Fix the backup — highest value overall

```bash
mysqldump -h "$DB_HOST" -u "$DB_USERNAME" …
```

…plus a **size assertion that fails the deploy** when the resulting file is implausibly small:

```bash
if [[ $(stat -c%s "$BK") -lt 1024 ]]; then
    log "FATAL: backup is ${?} bytes — refusing to migrate"
    exit 1
fi
```

A backup nobody verifies is not a backup. Fifty-six consecutive empty files were written, over two weeks, with every deploy reporting success. Do this even before anything else — until it is fixed, every future deploy still writes a 20-byte file.

### C. Remove the wildcard permission

Replace `Bash(php artisan *)` with a narrower allow list, and add explicit denials:

```json
"deny": [
  "Bash(php artisan migrate:fresh*)",
  "Bash(php artisan migrate:refresh*)",
  "Bash(php artisan migrate:reset*)",
  "Bash(php artisan db:wipe*)",
  "Bash(*DROP DATABASE*)",
  "Bash(*DROP TABLE*)"
]
```

Cost: more permission prompts during routine work. That is the correct trade for this box, where the same machine is both the workspace and production.

### D. A `PreToolUse` hook

Harness-level refusal of destructive database commands before execution — catching what A and C cannot: raw `mysql -e "DROP …"`, and env-var-prefixed invocations like the one that caused this.

### E. Least-privilege credentials *(larger change)*

The application's database user holds `DROP`. A separate migration-only account, with the runtime account lacking DDL rights, makes this class of accident impossible at the server rather than merely blocked at the client.

### F. Test-suite tripwire — **done**

`tests/Feature/DatabaseSafetyTest.php` asserts the suite targets a local, test-named database, and explains the cached-config mechanism in its docblock. The suite runs `RefreshDatabase`; if its isolation ever breaks, the failure mode is not a red test but a dropped production database.

---

## Operating rules that follow from this

1. **Never run `migrate:fresh`, `migrate:refresh`, `migrate:reset` or `db:wipe` from this checkout** — not with environment overrides, not with flags. The `.env` names production, and no prefix reliably changes that.
2. **Prove the target before any destructive command.** `php artisan db:show` prints the host and database actually in use. It costs two seconds and reads the same cached config the destructive command will.
3. **Run destructive commands in the foreground.** Backgrounding this one removed the moment where `Dropping all tables` could have been seen and interrupted.
4. **Treat an unexpected result as an alarm, not a puzzle.** The missing table in the scratch database was the signal. The first reading was "my migration didn't appear"; the correct reading was "then *where* did 449 migrations just run?"
5. **`config:cache` changes how every later command resolves its database.** Running it invalidates any assumption that `DB_*` overrides still work.

---

## Status

| Item | State |
|---|---|
| Database restored (mixed vintage) | ✅ partial — catalogue at 28 July |
| PropertySifu re-ingested | ✅ 301 |
| iProperty re-ingested | ✅ 151 |
| EdgeProp New Launches re-ingested | ✅ 2,735 — exactly the pre-wipe count |
| Missing permissions (`db:seed --class="\RolesSeeder" --force`) | ❌ outstanding |
| A — `prohibitDestructiveCommands` | ✅ done (`95178e45`) — verified by running `migrate:fresh --force` against production; it refused |
| B — backup `-h` + size assertion + fatal on failure | ✅ done (`95178e45`) — real dump is 188 MB; also fixed the `RELOAD`/`--single-transaction` fault this exposed |
| C — permission allowlist narrowed | ❌ outstanding |
| D — PreToolUse hook | ❌ outstanding |
| E — least-privilege DB user | ❌ outstanding |
| F — `DatabaseSafetyTest` | ✅ added |

---

*Prepared by Claude Code, which caused the incident. Written for the team that has to live with the consequences.*
