# Production setup — deploying (`scripts/deploy-update.sh`)

**Scope:** how a code update reaches production, and **how long the site is unavailable while it happens**. One command does the whole thing:

```bash
sudo bash scripts/deploy-update.sh
```

> Never `git pull` first. The script pulls, and it records the pre-deploy commit so `--rollback` has something to return to. Pulling by hand makes the working tree dirty, which the script (correctly) refuses to clobber.

## The maintenance window is SHORT by default

Until 2026-08-05 the script took the site down (`artisan down`) *before* the pull and lifted it *after* everything, so the 503 page covered the whole deploy — **2–5 minutes**, dominated by `npm run build` and the pre-migration `mysqldump`.

Now only the steps that genuinely need the site quiet are inside the window:

| Step | Site |
|---|---|
| `git fetch` / `reset --hard` | 🟢 live |
| `composer install` *(only if `composer.lock` moved)* | 🟢 live |
| `npm ci` *(only if `package-lock.json` moved)* | 🟢 live |
| `npm run build` — client + SSR bundles | 🟢 live |
| `wa-bridge` deps *(only if `wa-bridge/` changed)* | 🟢 live |
| `mysqldump` safety backup *(only if migrations changed)* | 🟢 live |
| `chown -R` + ACLs | 🟢 live |
| **`artisan down`** | |
| `migrate --force` | 🔴 **down** |
| `optimize:clear` → `config:cache` / `route:cache` / `view:cache` | 🔴 **down** |
| `systemctl reload php8.4-fpm` (clears OPcache) | 🔴 **down** |
| **`artisan up`** | |
| Horizon terminate, scheduler restart, Inertia SSR restart, wa-bridge *(if changed)* | 🟢 live |
| `service-check.sh` health check | 🟢 live |

**Typical downtime: 10–20 seconds.**

`mysqldump` reads a *live* InnoDB database happily, so nothing about the backup needed the site down. (It runs `--lock-tables=false`, **not** `--single-transaction` — the latter issues `FLUSH TABLES`, which needs the global `RELOAD` privilege the backup account doesn't hold; see [the 2026-08-07 incident](../../incident-2026-08-07-production-database-wipe.md). So it is a safety copy, not a consistent snapshot.) The worker restarts don't serve HTTP: Horizon and the scheduler are queue-side, and while the SSR daemon bounces, Inertia falls back to client-side rendering.

## What you trade for it

From the `git reset` until the window closes, the site is **live on new PHP source that does not yet have its new vendor packages, its rebuilt assets, or its migrations.**

- Pages this deploy did **not** touch are fine — the schema hasn't moved yet (migrate runs later, inside the window), so their unchanged queries still hold, and OPcache's default `revalidate_freq` bounds the raw source-swap mismatch to ~2s.
- A page the deploy **did** change can 500 for those couple of minutes, if it needs a column or a package that hasn't landed yet.

Note this is the **opposite** mismatch from an atomic-release (symlink) deploy, where *old* code briefly meets the *new* schema. Here the code moves first and the schema catches up — which is why **migrations do not have to be backward-compatible.** They run with the site down and the code already new, so a column drop or rename is as safe as it ever was.

What *does* stretch the window is a migration that rewrites a big table: it holds the lock for as long as it takes. **The window is only short if the migrations are.**

### The asset swap is atomic

Vite **empties its output directory** before writing the new bundle. Building straight into `public/build` therefore deletes `public/build/manifest.json` for the *whole* build — and the site is live throughout, because the window doesn't open until migrate, minutes later. `@vite()` in `app.blade.php` has nothing to read, so **every** request during the build 500s with `ViteManifestNotFoundException` — not just the pages the deploy touched, and the maintenance splash isn't up yet to hide it. That happened in production on 2026-08-12.

So the client bundle is built into a **staging directory** and swapped in with a rename:

1. `VITE_BUILD_OUT_DIR=public/build.new npm run build` — read by [`vite.config.js`](../../../vite.config.js) and applied to the **client** bundle only; the SSR bundle still lands in `bootstrap/ssr`. `laravel-vite-plugin` honours an explicit `build.outDir` while `base` keeps coming from its own `buildDirectory`, so the emitted URLs are `/build/...` wherever the files were written.
2. The outgoing build's hashed chunks are copied into `build.new` with `cp -an` — never clobbering, so the new `manifest.json` and new assets always win, while a tab loaded seconds earlier can still fetch a chunk it lazy-loads. `-a` preserves mtimes, which is what the prune keys on.
3. `mv build → build.old`, `mv build.new → build`, `rm -rf build.old`. Two renames on one filesystem: the gap where `public/build` doesn't exist is microseconds, not minutes.

Anything not rebuilt for 7 days is pruned, so `public/build` can't grow without bound. A **failed** build now also leaves the running site untouched, instead of stranding it on an emptied directory.

The same swap is packaged for the **dev loop** as [`scripts/live-build.sh`](../../../scripts/live-build.sh) — for a box that serves its working tree directly (Apache on `/var/www/html/peta`), where every small frontend edit needs a rebuild. It adds a `flock` so concurrent sessions queue, a `MemAvailable` guard, and skips itself when nothing is newer than the live manifest; it builds the client bundle only (`--ssr` opts in). Use it instead of `npm run build` or an `artisan down`/`up` wrapper — the latter merely turns the manifest-gap 500 into a 503 on every rebuild.

## Is anything half-done when the window opens?

**No.** The splash is served by `public/index.php` *before* Composer's autoloader — the framework never boots, so a request that lands in the window performs **no** write, queues **no** job, sends **no** mail. It is refused whole. The only cost is that the person has to repeat the action; there is never a partially-applied one to clean up.

A request that started *before* `artisan down` still finishes: `systemctl reload php-fpm` is graceful and waits for in-flight requests. (The one residual risk is such a request overlapping `migrate` and meeting a half-changed schema — a 10-20s exposure that has not bitten us.)

Sessions survive: they live on the `default` Redis connection while `cache:clear` only flushes the `cache` connection's own database — see the comments in [`config/database.php`](../../../config/database.php). **Nobody is logged out by a deploy.**

## What visitors see during the window

There are two paths, because a 503 means something different depending on what the person was doing.

**Full page load** (typing a URL, refresh, first visit) → [`resources/views/errors/maintenance.blade.php`](../../../resources/views/errors/maintenance.blade.php): a depleting brand-blue ring counting **10 → 0**, then **"Almost Done !!!"** with a spinner while it polls, and an automatic reload the moment the site answers with anything other than 503. Nobody has to press refresh — the page they wanted is what they get.

**Inertia XHR** (a form submit, a `<Link>` click) → [`resources/js/utils/maintenanceNotice.js`](../../../resources/js/utils/maintenanceNotice.js), a modal over the page they are already on. A reload is the wrong ending here: they are standing on a working page holding a half-filled form, and Inertia never navigated away, so **everything they typed is still there.** The notice therefore says the useful thing instead — *your last action was not saved, please do it again* — polls the same way, and on success turns green with a **Close and try again** button.

That path exists because Inertia's own error dialog cannot serve it. [`app.js`](../../../resources/js/app.js) cancels it for **503 only** (every other status keeps the dialog and its stack trace) via the cancelable `inertia:httpException` event. Inertia renders a non-Inertia response in an iframe with `sandbox="allow-scripts"` and **no** `allow-same-origin` — an opaque origin, where the splash's poll back to our own host is blocked by CORS and a reload would only reload the iframe. Dropped in there it spins on "Almost Done !!!" forever and the reader has to guess that Esc closes it. The blade detects that case (`window.top !== window.self`) and switches its own copy to *"Please try that again"*, so a tab still running a pre-2026-08-12 bundle isn't stranded either.

To see the modal without deploying anything, fire the event by hand in the browser console:

```js
document.dispatchEvent(new CustomEvent('inertia:httpException', {
    cancelable: true,
    detail: { response: { status: 503, headers: { 'retry-after': '15' } } },
}))
```

It is **pre-rendered to static HTML** by `artisan down --render="errors::maintenance"` into `storage/framework/maintenance.php`, and served by [`public/index.php`](../../../public/index.php) *before* Composer's autoloader and the framework boot. That ordering is the point: the window is exactly when the caches are being cleared and rebuilt, so booting the app to render an "we're down" page is when it's least likely to work.

Two consequences for anyone editing that view:

- **Only `$retryAfter` is available.** `DownCommand::prerenderView()` passes nothing else. Reading an undefined `$exception` is what used to make the render throw — the script then fell through to bare `artisan down`, and visitors got the plain 503 view saying *"check back in two hours"* for a 15-second deploy. `errors/503.blade.php` now accesses `$exception` null-safely so it can never re-break a prerender.
- **Everything must be inline.** No `@vite`, no `asset()`, no web fonts. `public/build` is being rewritten seconds before this page appears.

Three things still fall through to the framework rather than the splash, handled by Laravel's stub: paths in `PreventRequestsDuringMaintenance::$except`, the `--secret` bypass URL, and the bypass cookie.

> **Webhooks get 503 during the window.** The app has no custom `PreventRequestsDuringMaintenance`, so `$except` is empty — Stripe and Meta callbacks arriving in those ~15 seconds are refused. Both providers retry, so this has never needed fixing; if it ever does, add the middleware with an `$except` list and the stub honours it automatically.

## When to use `--full-maintenance`

```bash
sudo bash scripts/deploy-update.sh --full-maintenance
```

Restores the old behaviour — down for the entire deploy. Reach for it when:

- the deploy changes a **hot, high-traffic page** in a way that needs its new column or new package immediately;
- `composer.lock` moved a lot, or you're shipping something you're unsure of;
- you simply want nobody touching a half-updated tree.

## Other flags

| Flag | Effect |
|---|---|
| `--force-build` | Rebuild composer/npm deps even when the lockfiles didn't move |
| `--skip-pull` | Deploy the current working tree (no fetch/reset) |
| `--no-migrate` | Skip `migrate --force` |
| `--no-maintenance` | Never call `artisan down` at all |
| `--rollback` | Reset to the previous deploy's commit and rebuild (asks for confirmation; `--yes` skips) |
| `--branch <name>` | Deploy a branch other than the checked-out one |

Env overrides: `PHP_VERSION`, `APP_DIR`, `APP_USER`, `APP_GROUP`, `DEPLOY_BRANCH`, `AUTO_ROLLBACK`, `BACKUP_DB`.

Exit codes: `0` ok · `1` deploy failed (rolled back) · `2` bad args · `3` deployed but a service is unhealthy.

## When a deploy fails

The `ERR` trap **takes the site down first** — in short mode it was still serving a tree already known to be broken, and the recovery below re-runs composer and a full asset build, which is minutes — then resets the code to the previous commit, restores `vendor/`, rebuilds the assets, re-chowns, clears caches and reloads PHP-FPM, and finally lifts maintenance mode.

**Database migrations are never auto-reverted.** If one applied before the failure, restore a dump from `storage/app/deploy/backups/` (the last 10 are kept) or run `php8.4 artisan migrate:rollback` by hand.

Full timestamped log: `storage/logs/deploy.log`.

## Related files

| File | Role |
|---|---|
| [`scripts/deploy-update.sh`](../../../scripts/deploy-update.sh) | The deploy itself |
| [`resources/views/errors/maintenance.blade.php`](../../../resources/views/errors/maintenance.blade.php) | The countdown splash shown during the window |
| [`public/index.php`](../../../public/index.php) | Serves the pre-rendered splash before the framework boots |
| [`scripts/service-check.sh`](../../../scripts/service-check.sh) | Post-deploy health check (Apache, PHP-FPM, Horizon, scheduler, bridge) |
| [`scripts/server-setup.sh`](../../../scripts/server-setup.sh) | First-run provisioning — Apache vhost, PHP-FPM, systemd units, TLS |
| [`scripts/server-verify.sh`](../../../scripts/server-verify.sh) | Verifies a freshly provisioned box |
| [`horizon.md`](horizon.md) · [`scheduler.md`](scheduler.md) · [`reverb.md`](reverb.md) | The long-running services the deploy restarts |

## Not done here — true zero downtime

Getting to **0 seconds** means an atomic-release layout: build each deploy into `releases/<timestamp>/`, keep `.env` + `storage/` in a shared directory, and flip a `current` symlink at the end. That was scoped and deliberately deferred (2026-08-05) — it needs the Apache `DocumentRoot` moved, all four systemd units repointed, a one-off cutover with 1–3 minutes of downtime, and it reintroduces the backward-compatible-migration constraint this design avoids. The short window above is ~95% of the benefit for a fraction of the risk.
