# PropertyLab AI Evaluation — Operations Runbook

Audience: super admins operating the PropertyLab AI quality loop.
Scope: the evaluation subsystem shipped on branch `codex/propertylab-ai-evaluation`
(migrations, review UI, Cloud evidence evaluator, Evaluation dashboard,
GPU telemetry, guidance snapshots, lineage backfill).

## 1. Rollout

1. **Migrations** (all named-class, no schema-level foreign keys):
   - `2026_08_11_100000_create_zoom_meeting_analysis_reviews_table`
   - `2026_08_11_110000_add_unsupported_status_to_zoom_analysis_reviews`
   - `2026_08_11_120000_create_zoom_analysis_generations_table`
   - `2026_08_11_130000_create_zoom_analysis_evaluation_tables`
   - `2026_08_11_135000_create_propertylab_training_guidance_snapshots_table`
   - `2026_08_11_140000_create_peta_gpu_metric_samples_table`
   - `2026_08_11_141000_extend_peta_gpu_metric_samples`
   - `2026_08_11_142000_add_reliability_to_zoom_generation_attempts`
   - `2026_08_11_143000_harden_propertylab_evaluation_ledgers`
   Run `php artisan migrate` during a normal deploy window. Every `down()` is
   deterministic; the tri-state review migration preserves all legacy values
   on rollback (an `unsure` row keeps `has_hallucination = NULL`).
2. **Queue workers**: the evaluator runs on the `evaluation` queue
   (`redis-ai` connection) with its own single-process Horizon supervisor
   (`supervisor-evaluation`). Restart Horizon after deploy. Job-class changes
   need `sudo supervisorctl restart` of the worker, not `queue:restart`.
3. **Feature flags / env** (all default OFF or null — see `.env.example`):
   - `AI_EVALUATION_ENABLED` — Cloud evidence evaluation runs.
   - `AI_EVALUATION_DAILY_REQUESTS / _TOKENS / _COST_USD` — daily ceilings.
   - `AI_EVALUATION_MAX_PER_RUN / _MAX_ACTIVE_RUNS` — batch size and concurrent
     run caps (defaults 25 / 1; the active-run gate is database-serialized).
   - `AI_EVALUATION_MAX_UNITS / _MAX_OUTPUT_TOKENS / _MAX_INPUT_CHARS` and
     `AI_EVALUATION_MAX_REQUEST_COST_USD` — per-call bounds and declared
     worst-case reservation before spending.
   - `PETA_GPU_METRICS_ENABLED` (+ URL/token/limits) — private GPU telemetry.
   - `PETA_GPU_HOURLY_USD` — GPU price input; unset renders "Unavailable".
   - `PETA_GPU_FIXED_HOURLY_USD` — optional combined disk/EIP/network hourly
     input, displayed as its own duration-based estimate rather than merged
     into GPU or evaluator cost.
   - `PETA_AI_MODEL_REVISION / PETA_AI_IMAGE_DIGEST / APP_COMMIT` — immutable
     serving lineage. Missing values are preserved as null and set
     `lineage_missing=true`; they are never inferred from a friendly model name.

## 2. Dashboard access

`/manage/zoom/evaluation` — Zoom section tab "Evaluation". Requires `auth` +
admin + `view-zoom` permission + the `super-admin` middleware, and the
controller re-checks `isSuperAdmin()`. Ordinary admins receive 403; the tab is
hidden below super admin. Aggregate responses never contain transcripts, raw
AI request/response bodies, API keys, or reviewers' personal notes.
Evaluator `ai_requests` payloads are also hidden from ordinary admins and
cannot be replayed through the generic paid Compare action.

## 3. Evidence evaluation runs (Cloud, paid)

1. Preview first: **Preview run** on the dashboard (POST
   `/manage/zoom/evaluation-runs/preview`) — persists nothing, dispatches
   nothing, shows candidate identities + today's conservative reserved/actual
   budget usage, including cost.
2. Preview returns a 15-minute encrypted token that freezes actor, candidate
   IDs, period/limit, exact evaluator prompt and evaluator fingerprint.
   Changing the limit invalidates the browser preview. Confirm submits only
   this token; repeating it returns the same run and only recovers candidates
   missing a persisted queue-dispatch receipt.
3. Confirm: **Confirm run** (POST `/manage/zoom/evaluation-runs`). Creation
   fails BEFORE any dispatch when the flag is off, the configured
   conversation provider is `peta` (the evaluator never grades its own
   homework) or not allow-listed, or the daily budget is exhausted. Budgets
   are re-checked inside the job immediately before every Cloud call. A
   provider/model-priced worst-case token and cost reservation is persisted
   before network I/O; missing pricing, uncertain retries and oversized
   system+user prompts fail closed. Work
   skipped there settles `skipped_budget` and the run ends
   `completed_with_errors`, never silent success.
4. Candidates are frozen at creation (hard cap `AI_EVALUATION_MAX_PER_RUN`,
   default 25) and each has a one-time settlement ledger; empty/deleted
   generations, post-claim crash recovery and queue redelivery cannot increment
   counters twice. A generation without exact transcript hash lineage is
   skipped before Cloud evaluation.
   By default only one run may be queued/running (`AI_EVALUATION_MAX_ACTIVE_RUNS=1`).
5. Retry/failure: transient provider failures retry with AiJob backoff; a
   terminal failure settles the candidate's pending units as `failed` and the
   run's counters stay truthful. A wholly failed run is `failed`, a mixed one
   `completed_with_errors`.

**A real 10-meeting pilot spends real vendor money — run it only inside a
cost window the business owner has explicitly confirmed.** Automated tests
never call the paid evaluator.

## 4. Metrics, formulas and cost labels

Every dashboard figure uses the frozen formulas in the design (§6) via
`PropertyLabEvaluationMetrics`. Non-negotiable semantics:

- Missing data renders as **unavailable/null — never zero**.
- `Not sure` and unanswered are **excluded** from the confirmed-issue
  denominator and shown separately.
- **Automatic evidence flags** (transcript-verifiable units only, diagnostic)
  and **human-confirmed issues** are separate figures, and neither is ever
  labelled a "hallucination rate".
- P95 is nearest-rank (never interpolated) and marked unstable under 20
  samples.
- Costs stay separate: Cloud evaluator vendor cost (**Estimate**, logged token
  usage × configured price table), **PropertyLab GPU cost
  (always labelled Estimate — duration-derived, may double-count
  concurrency)**, and benchmark allocated cost (only a verified billable run
  window may carry this label). Optional disk/EIP/network hourly cost is a
  fourth, separately labelled estimate. **No combined "total AI cost" exists.**

## 5. GPU telemetry

- Server-side pull of ONE fixed private vLLM `/metrics` URL —
  **no browser-to-vLLM call, no Laravel-to-GPU SSH, ever**. The URL comes
  only from server config, redirects are refused, responses are bounded, and
  only allow-listed numeric series are stored (never payloads/labels).
- Default OFF. With the flag off the scheduler registers nothing — zero
  per-minute traffic. A permanently `unavailable` state is an acceptable
  verified result while no private route exists.
- Staleness: a sample older than `PETA_GPU_METRICS_STALE_AFTER` (default
  180 s) renders as **Stale**; no sample renders as **Unavailable** ("no
  data", explicitly not zero load).
- Memory stability uses the three newest comparable verified benchmark
  windows with explicit idle samples (`requests_running=0` and
  `requests_waiting=0`) before and after each run. Busy or missing samples
  remain `insufficient`; they never support a "stable" claim.
- The latest card and stability windows use the configured source+instance.
  OOM/restart counter resets are warnings, never clean zeroes.
- GPU infrastructure changes and ECS lifecycle actions (start/stop/resize/
  release, security groups, port exposure) remain manual, human-confirmed
  operations outside this application.

## 6. Guidance and version changes

- Guidance is `propertylab-guidance-v1` — deterministic, advisory, and
  auditable (rule version, sample, thresholds, case IDs on every card).
- **Ten meetings with 50 % coverage validate the collection UX only.** Pilot
  data shows `pilot_signal` cards at most; Prompt/RAG/fine-tune each require
  their frozen minimum evidence (30 completed generations + 100 confirmed
  units to qualify at all; RAG additionally 30 external-knowledge errors at
  ≥ 60 % of actionable errors + a declared corpus; fine-tuning additionally
  500 adjudicated development + 100 held-out examples, ≥ 99 % schema, a
  ≥ 10 % persistent error across ≥ 2 prompt versions). Holdout rows are never
  training candidates; cohorts are deterministically meeting-grouped at
  generation/backfill time. Until a dedicated adjudicator workflow exists,
  an example qualifies conservatively only when at least two human reviewers
  give the same definitive verdict (and the same issue category for a
  problem); conflicts and `unsure` stay outside training counts.
- Prompt text changes rotate the stored prompt SHA automatically; schema
  changes must bump `ConversationAnalysis::SCHEMA_VERSION` (a frozen unit
  test enforces this); rule changes must ship as a NEW guidance version, and
  historical snapshots keep the old one.

## 7. Lineage backfill (one-off)

```
php artisan zoom:backfill-analysis-generations --limit=200 --dry-run
php artisan zoom:backfill-analysis-generations --limit=200
```

`--limit` is required (capped at 500). Idempotent, PropertyLab-only, never
calls AI, never dispatches evaluation, no schedule. Reconstructed rows carry
`lineage_source=backfill` and first-pass fields stay NULL (unknown, not
success). Rows with an online generation still processing are skipped, so
backfill cannot claim that generation's request or projection.

## 8. Privacy and export boundaries

Review units, evidence quotes and rationales are bounded excerpts;
transcripts stay on the recording detail behind existing visibility rules.
Per-user claim opinions are private to each reviewer in the UI; only
aggregate counts reach the dashboard. There is no bulk export surface in this
phase — adding one requires its own review of these boundaries.

## 9. Rollback

- Feature-level: turn the flags off (`AI_EVALUATION_ENABLED`,
  `PETA_GPU_METRICS_ENABLED`) — everything else keeps working.
- Review UI: the legacy boolean is dual-written; rolling back the tri-state
  migration restores it deterministically for every row.
- Schema-level: reverse the migrations newest-first. Generation history and
  evaluation runs are append-only — dropping them loses history but never
  corrupts the latest projections (`zoom_meeting_analyses` remains the UI's
  source of truth).
