# System Health (Manage)

**Portal:** Manage · **Route:** `manage.system.health` (`/manage/system/health`) · **Nav:** Others → **System → System Health** (Operations suite; the tab is dropped from the strip for non-super-admins) · **Access:** **super-admin ONLY** (`role:super-admin` on the routes + the nav item hidden for everyone else — raw logs can contain message bodies and tokens)

## What it does
One page that answers "is everything actually running?" without SSH: live status cards for the **Scheduler** (schedule:work), **Horizon** (queue workers + per-supervisor lanes), **Reverb** (WebSocket), the **Baileys wa-bridge** sidecar (process + per-instance socket states + per-channel `last_ping_at` + webhook-DLQ depth), **Redis**, the **Database**, **failed jobs / queue backlog**, and **Payments** (whether any gateway is enabled and fully configured) — plus the **upcoming schedule** (every `Kernel::schedule()` entry with its cron expression and next due time, the browser equivalent of `php artisan schedule:list`) and a **storage/logs viewer** (tail any log file, parsed and colorized, incl. `wa-bridge.log`).

## How it works
- **`SystemHealthService`** ([src/Common/Services/SystemHealthService.php](/src/Common/Services/SystemHealthService.php)) — every probe is defensive (reports `ok | warn | down | off` + a human detail line, never throws):
  - **Scheduler** — reads the `ops:scheduler-heartbeat` cache stamp that the FIRST entry in [app/Console/Kernel.php](/app/Console/Kernel.php)'s schedule writes every minute. Stale >3 min ⇒ `schedule:work` is down ⇒ EVERY scheduled automation (WhatsApp heartbeat, broadcasts, flow reaper, …) is silently stopped. This is the check that catches "I deployed but forgot the scheduler" — locally it means `php artisan schedule:work` isn't running in WSL.
  - **Horizon** — Horizon's own `MasterSupervisorRepository` / `SupervisorRepository` (masters running/paused + each supervisor's queues and process counts).
  - **Reverb** — a raw 2s TCP connect to `broadcasting.connections.reverb.options.host:port`. Reports `off` (not `down`) when `WHATSAPP_REALTIME` is disabled — the inbox intentionally runs on the 6s poll then.
  - **Bridge** — `BridgeGateway::health()` (the wa-bridge `GET /health`, now enriched with a per-instance `{name, state, last_ping_at, reconnect_attempts, conflict}` summary + the webhook-DLQ row count) + every Bridge channel's DB `status` / `last_ping_at` freshness (stamped by `whatsapp:ping-bridge`).
  - **Redis / Database** — `PING` / `SELECT 1` with latency.
  - **Jobs** — `failed_jobs` total + last-24h, and Horizon's per-queue workload (length + wait).
  - **Upcoming schedule** — resolves a fresh `Schedule` and invokes `Kernel::schedule()` on it via reflection (the Schedule is only populated during console bootstrap), then maps each event to `{command, expression, next_due}` sorted by next-due.
- **`LogViewerService`** ([src/Common/Services/LogViewerService.php](/src/Common/Services/LogViewerService.php)) — lists `storage/logs/*.log` (+ rotated `.log.N`); tails the last N lines of one by **seeking backward from the end** (capped 5 MB — a huge file is never loaded whole); parses the tail: `wa-bridge*.log` as **pino JSON** (numeric levels → labels; `time` is ISO-8601 on new lines, epoch-ms on old ones — both handled), `laravel*.log` as `[timestamp] env.LEVEL: message` entries with stack-trace continuation lines folded into their entry, anything else raw. File access is locked to basenames matching `*.log(.N)` inside `storage/logs` (no traversal). Default file = today's `laravel-Y-m-d.log`, else the newest.
- **`HealthController`** ([app/Http/Controllers/Manage/System/HealthController.php](/app/Http/Controllers/Manage/System/HealthController.php)) — `index` (Inertia page with the initial snapshot), `data` (JSON refresh, polled), `logs` (file list JSON), `logTail` (parsed tail JSON, `?file=&lines=`, lines clamped 10–2000).
- **`Health.vue`** ([resources/js/Pages/Manage/System/Health.vue](/resources/js/Pages/Manage/System/Health.vue)) — two `ShowTabs` tabs:
  - **Services** — status cards (with the Horizon supervisor lanes, queue workload, Bridge channels/instances/DLQ inline) + the upcoming-schedule table, **grouped by module** (whatsapp / ops / calls / zoom — the server derives `group` from the command prefix) with a **human frequency label** ("Every minute" / "Every 5 minutes" / "Daily at 03:00", the raw cron kept as small print). Auto-refreshes every 15s via a silent axios poll of `health/data` (paused while the tab is hidden; toggle + manual Refresh). ⚠️ `upcomingSchedule()` must resolve the console Kernel BEFORE reading the Schedule singleton and only reflection-invoke `schedule()` when the event list is still empty — the Kernel's booted hook already populates it in HTTP context, and invoking unconditionally double-registers every event (the duplicated-rows bug).
  - **Logs** — file list grouped **Today-first** (files named with today's date, or written today — the legacy un-dated `wa-bridge.log` during the transition) with older days collapsed behind an "Older days (N)" expander (auto-expanded only when nothing was logged today) · viewer with lines selector (100–1000), level filter chips (parsed formats), client-side text filter, a 5s **Live** auto-tail toggle, and colorized entries (time · level badge · message · dimmed context/stack).
- **Ops alerting (related):** [src/Common/Services/OpsAlertService.php](/src/Common/Services/OpsAlertService.php) + [config/ops.php](/config/ops.php) — keyed, throttled alerts (log always; mail via `OPS_ALERT_MAIL_TO`; Slack-compatible webhook via `OPS_ALERT_WEBHOOK_URL`). Wired into [whatsapp:ping-bridge](/app/Console/Commands/PingWhatsappBridge.php): a Bridge channel's CONNECTED→DISCONNECTED transition, its recovery, and "every ping failed AND /health is dead" (bridge process down) all alert.
- **Log alert digest (`ops:scan-logs`, every 5 min):** [src/Common/Services/LogAlertScanner.php](/src/Common/Services/LogAlertScanner.php) + [app/Console/Commands/ScanLogsForAlerts.php](/app/Console/Commands/ScanLogsForAlerts.php) — scans the lines APPENDED to today's + yesterday's `laravel-*.log` **and `wa-bridge-*.log`** (plus the legacy un-dated `wa-bridge.log` until the bridge restarts onto the daily scheme) since the previous run (per-file byte **watermarks** in `storage/framework/log-alert-watermarks.json` — a JSON file, deliberately NOT cache, so `cache:clear` never resets it; first sighting of a file initialises at EOF so history is never flooded; a shrunken file = rotation → rescan from 0; ≤2 MB read per file per run) and emails **one plain-text digest per run** of every entry at/above `OPS_LOG_ALERT_LEVEL` (default `warning`) to `OPS_LOG_ALERT_MAIL_TO` (comma-separated; falls back to `OPS_ALERT_MAIL_TO`; the schedule entry is `->when()`-skipped when neither is set). Levels: laravel `env.LEVEL:` header lines (stack-trace continuations are skipped — the digest carries the headline, the viewer has the trace); wa-bridge pino numeric levels (≥40 warn / ≥50 error). Noise control: `OPS_LOG_ALERT_IGNORE` (|-separated substrings), digest capped at `OPS_LOG_ALERT_MAX_ENTRIES` (default 50, "+N more"), and **notify-once repeat suppression** — see below.
- **Notify ONCE, then hold (`OPS_LOG_ALERT_REPEAT_AFTER`, default 1440 min).** The watermark only promises each *line* is reported once, which is no help when the broken thing writes a **new** line every few minutes — the real case being a **blocked WhatsApp number**: `whatsapp:ping-bridge` runs every minute, sees the instance parked `dead`, revives it with a fresh budget, the budget burns, and the bridge logs `gave up after N failed reconnects — parked as dead` again, forever (~12 min apart in practice). The digest then mails the same sentence all day and the recipient stops reading it, which is how the one alert that mattered gets missed. So a message is emailed once and **held**; occurrences during the hold are **counted, not discarded**, and the digest that carries it again appends `(+N more since …)`. Two lines are "the same message" when they match after **digits are collapsed** — the retry count is the varying part, and a literal comparison would call every retry a new alert; everything else (the instance name, the file) is kept, so one broken number never mutes another. Two identical lines inside ONE run fold together too. State lives in `storage/framework/log-alert-repeats.json`, a sibling of the watermark file and a plain file for the same reason (`cache:clear` must not turn every held message back into an email), committed **with** the watermarks so an undelivered digest is never silently held. `0` restores the old every-occurrence behaviour. Prefer this over `OPS_LOG_ALERT_IGNORE`, which hides a message permanently — including on the day it starts meaning something new. Covered by [tests/Feature/System/LogAlertRepeatTest.php](/tests/Feature/System/LogAlertRepeatTest.php). **Self-feed safe:** a digest-send failure is logged at *warning*… which the next run would re-report once, but the mail MIME the `log` driver writes is DEBUG + continuation lines, so a delivered digest never re-triggers itself. Test with `php artisan ops:scan-logs --dry-run` (detect + print, no email, watermarks untouched).

## Related files
- [src/Common/Services/SystemHealthService.php](/src/Common/Services/SystemHealthService.php) — the probes + upcoming schedule.
- [src/Common/Services/LogViewerService.php](/src/Common/Services/LogViewerService.php) — log list / tail / parse.
- [src/Common/Services/OpsAlertService.php](/src/Common/Services/OpsAlertService.php) · [config/ops.php](/config/ops.php) — keyed + throttled ops alerts (mail / webhook, opt-in by env).
- [src/Common/Services/LogAlertScanner.php](/src/Common/Services/LogAlertScanner.php) · [app/Console/Commands/ScanLogsForAlerts.php](/app/Console/Commands/ScanLogsForAlerts.php) — the `ops:scan-logs` watermarked error/warn digest email (watermarks in `storage/framework/log-alert-watermarks.json`, repeat holds in `log-alert-repeats.json`; both committed only after a successful send, so a mail outage delays the digest instead of dropping it).
- [tests/Feature/System/LogAlertRepeatTest.php](/tests/Feature/System/LogAlertRepeatTest.php) — the notify-once hold: emailed once then held, the hold expiring with a `(+N more)` count, digits collapsing so a changing retry count is still one problem, a different message never muted by another's hold, same-run duplicates folding, `0` restoring every-occurrence, and an undelivered digest not marking a message as notified.
- [app/Http/Controllers/Manage/System/HealthController.php](/app/Http/Controllers/Manage/System/HealthController.php) — index / data / logs / logTail.
- [resources/js/Pages/Manage/System/Health.vue](/resources/js/Pages/Manage/System/Health.vue) — the page (Services + Logs tabs).
- [app/Console/Kernel.php](/app/Console/Kernel.php) — the `ops-scheduler-heartbeat` schedule entry (first in the list).
- [routes/web.php](/routes/web.php) — `manage.system.*` (standalone group, `['auth','admin','role:super-admin']`).
- [resources/js/Layouts/ManageLayout.vue](/resources/js/Layouts/ManageLayout.vue) — the "System Health" nav item (Operations suite → Others, super-admin only).
- [scripts/service-check.sh](/scripts/service-check.sh) — the SSH-side equivalent check (php-fpm/apache/redis/Horizon/scheduler/bridge).

## Gotchas
- **The scheduler card is self-referential in one direction only:** the heartbeat is written BY the scheduler, so "ok" proves it runs; but if the page itself is unreachable nobody sees it — pair with an external uptime check on the app for full coverage.
- The upcoming-schedule table shows when things WOULD fire; it does not prove they fired — the Scheduler card's heartbeat does that.
- wa-bridge log lines written before 2026-07-08 carry epoch-ms `time` values; the viewer converts both formats to the app timezone.
- Log tail is read-only and capped (2000 lines / 5 MB back-read); it is a triage tool, not a full log browser.
- **A repeating `gave up after N failed reconnects — parked as dead` is a symptom, not noise.** The log line tells the reader to "re-open the QR modal to retry", but [`ping()`](/wa-bridge/src/instances.js) does not wait for that: the every-minute heartbeat auto-revives any instance in state `dead` with a fresh reconnect budget. That is deliberate and right for a transient outage, and wrong for a permanent one — a **blocked or banned number** can never connect, so the bridge reconnects against it around the clock. `OPS_LOG_ALERT_REPEAT_AFTER` quiets the mail; it does not stop the reconnect attempts. There is no backoff on the revive path yet.
