# Production setup — Horizon (queue workers)

**Scope:** how to run **Laravel Horizon** (the queue worker pool) in production, and how to **deploy new code with zero downtime**. Horizon is a long-running process, so it does NOT pick up new code on its own — see *Deploying* below (this is the #1 footgun).

> Everything that "does work" in this app rides the queue: inbound webhooks, sending WhatsApp messages, AI replies, read receipts, and **broadcast events** (real-time). **No running Horizon ⇒ none of that happens** (the inbox falls back to a 6 s poll, but nothing is sent/processed). See also [realtime.md](/docs/modules_handbook/manage/messages/whatsapp/realtime.md) and [reverb.md](/docs/modules_handbook/production-setup/reverb.md).

## What must be running

| Piece | Why | Process manager |
|---|---|---|
| **Redis** | the queue backbone + Horizon's own metadata | system service (apt/systemd) |
| **`php artisan horizon`** | the worker pool (one master → its supervisors) | Supervisor / systemd — **exactly one** |
| (web) PHP-FPM + Nginx | serves the app + dispatches jobs | — |

Plus two sibling daemons documented separately: **Reverb** ([reverb.md](/docs/modules_handbook/production-setup/reverb.md)) and the **wa-bridge** Node sidecar ([wa-bridge/README.md](/wa-bridge/README.md)) — both also need Supervisor.

## Required environment (`.env`)

```dotenv
APP_ENV=production            # REQUIRED — Horizon reads the `production` block in config/horizon.php
                             # (supervisor-1 → 10 procs, supervisor-ai → 4). Without it you get the
                             # `defaults` (1 proc each) and the app can't keep up.
QUEUE_DRIVER=redis            # config/queue.php reads QUEUE_DRIVER (NOT QUEUE_CONNECTION); default is
                             # `sync` — if this is unset/wrong, jobs run inline on the web request.
REDIS_HOST=127.0.0.1
REDIS_PORT=6379
REDIS_PASSWORD=…             # set a real password in production
```

> Cache & sessions are **file**-based here (`CACHE_DRIVER=file`, `SESSION_DRIVER=file`), so Redis is used **only** for the queue + Horizon. Set Redis `maxmemory-policy` to **`noeviction`** (or give it enough memory) — an eviction policy that drops keys can silently lose queued jobs.

## The lanes (from `config/horizon.php` + `config/queue.php`)

Two isolated supervisors so a slow/down AI provider can never starve normal sends:

| Supervisor | Connection / queue | Prod procs | Job timeout | Queue `retry_after` |
|---|---|---|---|---|
| **supervisor-1** | `redis` / `default` | up to **10** | 180 s | 190 s |
| **supervisor-ai** | `redis-ai` / `ai` | up to **4** | 320 s | 390 s |

The invariant to preserve if you ever tune these: **job `timeout` < queue `retry_after`** (180<190, 320<390) — otherwise the queue re-delivers a job that is still running and it executes twice (a double WhatsApp send, or a double-billed AI call). `App\Jobs\Ai\AiJob` subclasses ride the `ai` lane; everything else rides `default`.

## Run it with Supervisor (recommended)

`/etc/supervisor/conf.d/horizon.conf`:

```ini
[program:horizon]
process_name=%(program_name)s
command=php /var/www/petav3/artisan horizon
directory=/var/www/petav3
autostart=true
autorestart=true
user=www-data            ; the deploy user that owns the code + storage
redirect_stderr=true
stdout_logfile=/var/www/petav3/storage/logs/horizon.log
stopwaitsecs=3600        ; ≥ the longest job (AI lane up to 320 s) so a restart never SIGKILLs a job mid-run
numprocs=1               ; CRITICAL: exactly ONE horizon master — see "Only one master" below
```

```bash
sudo supervisorctl reread && sudo supervisorctl update
sudo supervisorctl start horizon
```

> On **Laravel Forge**: enable the project's **Horizon** toggle (Forge writes this exact Supervisor entry and keeps it alive) — don't also start a manual `php artisan horizon`, or you get two masters.

## The scheduler (cron) — separate from Horizon

Horizon runs *queued* jobs; *scheduled* commands (e.g. `calls:poll-dowayai`) need the Laravel scheduler, which needs one cron line:

```cron
* * * * * cd /var/www/petav3 && php artisan schedule:run >> /dev/null 2>&1
```

Optional: add `$schedule->command('horizon:snapshot')->everyFiveMinutes();` in `app/Console/Kernel.php` to populate the Horizon **metrics** graph (cosmetic — not required for jobs to run).

## Deploying — update code with (near) zero downtime

Horizon workers are **long-running PHP processes that hold compiled code in memory** — they keep running the OLD code until restarted. After every deploy that touches `app/`, `src/`, `config/`, `routes/`, or `composer`, you MUST tell Horizon to restart. The order:

```bash
cd /var/www/petav3
git pull                                   # or Forge/Envoyer release
composer install --no-dev --optimize-autoloader
npm ci && npm run build                    # frontend assets (also bakes VITE_* — see reverb.md)
php artisan migrate --force                # DB changes (e.g. the read-tracking migration)
php artisan config:cache && php artisan route:cache && php artisan view:cache
php artisan horizon:terminate              # ← graceful restart: see below
php artisan reverb:restart                 # ← restart the WebSocket server too (see reverb.md)
```

**Why `horizon:terminate` is the seamless way:** it tells the master to **finish the jobs currently in flight, then exit**. Because `fast_termination=false` (config/horizon.php), it waits for running workers (no job is killed mid-run). Supervisor sees the master exit and — because `autorestart=true` — immediately starts a **fresh master that loads the new code**. In-flight jobs complete on the old code; everything dispatched after restart runs the new code. There is no manual stop/start and no lost jobs.

> **Always run `horizon:terminate` after deploying.** The classic bug (we hit it): you change a Job's code, but the running workers keep executing the *old* version, so the new behaviour silently never takes effect — or worse, orphaned old workers keep draining the queue. If `config:cache` ran, terminate *after* it so workers boot with the fresh cached config.

## Only one master (avoid orphan workers)

`php artisan horizon` is a **master** that spawns the per-supervisor worker children. If a master dies without cleanly stopping its children (e.g. `Ctrl-C` in a stray terminal, or starting Horizon twice), the **child workers are orphaned and keep pulling jobs** — so jobs get processed by an invisible old process and your real Horizon shows nothing. Rules:

- **Exactly one** master (`numprocs=1`; never a manual `php artisan horizon` on top of the Supervisor one).
- To stop everything cleanly: `php artisan horizon:terminate` (graceful) — not `Ctrl-C`.
- If you suspect orphans: `ps aux | grep '[h]orizon'` should show **one** `horizon` master + its `horizon:work …` children all sharing the **same** `--supervisor=` id. Multiple supervisor ids = leftover masters → `pkill -f 'artisan horizon'` then start one.

## Monitoring & security

- **Dashboard:** `/horizon` — already locked down by the `viewHorizon` gate in `app/Providers/HorizonServiceProvider.php` (admins only, `$user->isAdmin()`). Keep it behind `auth`.
- **Wait alerts (`config/horizon.php` → `waits`):** `redis:default` fires `LongWaitDetected` at **60 s**, `redis-ai:ai` at **300 s** (AI is expected to be slow). Wire a notification channel if you want paging.
- **Failures:** the `failed_jobs` table (Horizon → *Failed Jobs*); retry from the dashboard or `php artisan queue:retry all`.
- **Health:** `php artisan horizon:status` (running / paused). Memory: the master is capped at 64 MB (`memory_limit`), each worker 128 MB (`memory`) — it self-restarts on breach.

## Quick triage

| Symptom | Likely cause → fix |
|---|---|
| Nothing processes; jobs pile up | Horizon not running (`supervisorctl status horizon`) or Redis down (`redis-cli ping`). |
| New code/behaviour not taking effect | Workers run stale code → `php artisan horizon:terminate` (Supervisor restarts fresh). |
| Jobs run but the real Horizon terminal is empty | **Orphan masters** draining the queue → `pkill -f 'artisan horizon'`, start one. |
| AI replies fail / time out, `default` jobs fine | The `ai` lane is isolated by design; check the provider/network, not Horizon (see [shared/ai/readMe.md](/docs/modules_handbook/shared/ai/readMe.md) Resilience). |
| Real-time dead but sends work | Broadcasts ride the queue → Horizon is fine; **Reverb** is down → see [reverb.md](/docs/modules_handbook/production-setup/reverb.md). |

## Related files
- [config/horizon.php](/config/horizon.php) — supervisors, prod scaling, wait alerts, trim, `fast_termination`.
- [config/queue.php](/config/queue.php) — `redis` / `redis-ai` connections + the `retry_after` invariant.
- [app/Providers/HorizonServiceProvider.php](/app/Providers/HorizonServiceProvider.php) — the `/horizon` admin gate.
- [app/Console/Kernel.php](/app/Console/Kernel.php) — the scheduler (needs the `schedule:run` cron).
- [reverb.md](/docs/modules_handbook/production-setup/reverb.md) — the real-time WebSocket sibling · [realtime.md](/docs/modules_handbook/manage/messages/whatsapp/realtime.md) — how broadcasts flow.
