# petav3-dev Incident Summary — Memory Exhaustion Causing HTTP 522

**Date:** 2026-08-02
**Host:** `petav3-dev` (GCP VM, asia-southeast1-b, external IP 35.213.165.89)
**Impact:** `https://wk.propertylab.com.my` intermittently unreachable (Cloudflare Error 522),
SSH access also became unresponsive on at least one occasion.

## Context: this box is production, not a sandbox

Despite the hostname `petav3-dev`, this VM has `APP_ENV=production` in `.env` and serves
the real live domain `https://wk.propertylab.com.my`. There is no separate staging/dev
environment — this machine does double duty as both the dev workspace (Claude Code, VS Code
Remote) and the production web server. That combination is the root of the incident.

## Root cause

The VM has **4GB RAM and no swap configured at the OS level initially** (swap turned out to
already exist at 2GB, likely set up previously, but wasn't being effectively relied upon).

Baseline idle memory usage is reasonable (~1.3GB used) from:
- Apache (11 workers)
- PHP-FPM
- Laravel Horizon (6 supervisors + 6 workers across queues: default, ai, broadcast, f2f,
  transcription, video)
- Node.js app server (`src/server.js`)
- Google Cloud Ops Agent (~200MB)

**The actual trigger:** a Claude Code session, while working on a frontend task, autonomously
ran `npm run build` (`vite build && vite build --ssr`) to verify its changes. This single
process spiked to **~1.17GB RAM (29% of total) and ~19% CPU**, combined with 3 concurrent
Claude Code sessions + VS Code Remote Server + extension host already running
(collectively another ~1GB+). Total demand exceeded available RAM, driving `available` memory
from ~2.5GB down to ~27MB within about 10 minutes.

Once memory was exhausted, load average spiked to **29.73** (1-min avg), and the box became
unable to service requests — even `curl localhost` timed out locally. In the worse instance,
this also made the network stack unresponsive enough that SSH (both browser-based and native
client) timed out completely, requiring a GCP **Reset** (hard reboot) to recover.

## What we ruled out

- **Disk space:** fine (66-68% used, plenty of headroom).
- **Apache/PHP-FPM crash:** service was always `active (running)`; the problem was resource
  starvation, not a crashed service.
- **GCP Firewall:** `default-allow-ssh` rule (tcp:22, Allow, apply to all) was correctly
  configured — not a network ACL issue.
- **Memory leak (single process growing unbounded over time):** not confirmed — the spike was
  sudden and tied directly to the `vite build` invocation, not a slow creeping leak. (A
  background monitoring loop logging `free -h` + `ps aux --sort=-%mem` every 15 min to
  `~/memory_log.txt` was set up to watch for this longer-term, but the acute cause was found
  before it accumulated much data.)

## Fixes applied

1. **Killed the runaway `vite build` process** (`kill -9 <pid>`) — immediately freed ~1GB,
   memory `available` went from 27Mi → 1.0Gi, site returned `200 OK` within seconds.
2. **Confirmed swap is present** (2GB) and actively being used as a buffer — this should help
   prevent a full network-stack lockup next time, though it doesn't eliminate the underlying
   resource contention.
3. Verified no orphaned `vite`/`npm run build` processes remained after the incident.

## Key operational learning

**This is not primarily a "4GB RAM is too small for production" problem.** Idle/steady-state
load fits comfortably in 4GB. The real problem is **running heavyweight, bursty dev tooling
(frontend builds, multiple concurrent Claude Code agent sessions, VS Code Remote) on the same
box that serves live production traffic.** Any one of these can transiently spike RAM/CPU
enough to starve Apache, regardless of the "steady state" baseline.

## Standing SOP for next time the site is unreachable

SSH in and run this one-liner first:

```bash
echo "=== TIME ===" ; date ; echo "" ; echo "=== MEMORY ===" ; free -h ; echo "" ; \
echo "=== LOAD AVERAGE ===" ; uptime ; echo "" ; \
echo "=== TOP MEMORY PROCESSES ===" ; ps aux --sort=-%mem | head -15 ; echo "" ; \
echo "=== TOP CPU PROCESSES ===" ; ps aux --sort=-%cpu | head -10 ; echo "" ; \
echo "=== APACHE STATUS ===" ; sudo systemctl status apache2 --no-pager ; echo "" ; \
echo "=== LOCAL CURL TEST ===" ; time curl -I -m 10 -H "Host: wk.propertylab.com.my" http://localhost ; echo "" ; \
echo "=== DISK SPACE ===" ; df -h
```

- If a specific process is hogging memory/CPU (as with `vite build`), `kill -9 <pid>` it and
  re-test with `curl`.
- If SSH itself won't connect (network stack unresponsive), the only recovery path found so
  far is **GCP Console → VM instances → petav3-dev → Reset** (hard reboot, ~1-2 min downtime,
  disk/data untouched, safe given static IP is configured).
- Background monitor (optional, useful for spotting slow leaks over time):
  ```bash
  nohup bash -c 'while true; do date; free -h; echo "---PROCESSES---"; ps aux --sort=-%mem | head -15; echo "==========="; sleep 900; done' >> ~/memory_log.txt 2>&1 &
  disown
  ```
  Check later with `tail -100 ~/memory_log.txt`.

## Open questions / things to decide with the team

1. **Should Claude Code (and any agentic dev tooling) be moved off this box entirely?**
   Given it autonomously ran a production-impacting build without knowing the resource
   constraints, this is the most direct fix.
2. **Should frontend builds be moved to CI/CD or done locally**, with only the compiled
   `public/build/` output deployed to this server — so this box never runs `vite build`
   directly at all?
3. **Is a real staging/dev environment (separate VM) worth provisioning**, so dev work
   (including Claude Code sessions) never shares resources with the box serving live traffic?
4. Cost note: bumping this VM's machine type (e.g. `e2-medium` 4GB → `e2-standard-2` 8GB) was
   considered, but given the root cause is workload placement rather than baseline capacity,
   the team should decide whether that spend is worth it versus fixing tooling placement.

## Rule proposed for `CLAUDE.local.md` (not yet added — pending discussion)

```markdown
## Resource constraints — critical

- This box has only 4GB RAM and serves live production traffic. Do NOT run `npm run build`,
  `vite build`, or `npm run dev` ad-hoc to verify frontend changes — it can consume 1GB+ RAM
  and CPU spike enough to make Apache stop responding (has happened repeatedly).
- To verify frontend changes, read the code and reason about correctness instead of building.
  If a build is genuinely necessary, ask the user first and warn it may cause a brief outage.
- Never run build/deploy commands automatically as part of a coding task unless the user
  explicitly asked to deploy via `/sync`.
```
