# PropertyLab AI Evaluation and Training Guidance — Implementation Plan

> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development
> to implement this plan task-by-task. Use a fresh implementation worker for each task, followed by an independent task review. Do not run implementation workers in parallel.

**Goal:** Deliver the approved V1+V2 PropertyLab AI quality system: business-friendly human reviews, immutable generation lineage, Cloud AI evidence checks, per-user claim confirmation, super-admin metrics/guidance, and safe GPU telemetry.

**Architecture:** Keep `zoom_meeting_analyses` as the latest UI projection and add append-only generation/evaluation records for history. Extend the existing `AiClient`/`AiJob` path rather than creating a second provider stack. Automatic evidence verdicts are immutable suggestions; separate per-user review records preserve human judgment. A protected Zoom Evaluation page aggregates explicit formulas and calls a pure deterministic guidance service. GPU telemetry is pulled server-side from one fixed private endpoint and fails closed/unavailable.

**Tech stack:** Laravel, Eloquent, MySQL, Inertia, Vue 3, Tailwind, PHPUnit, Vitest/happy-dom, Laravel queues/Horizon, OpenAI-compatible vLLM metrics.

**Approved design:** `docs/superpowers/specs/2026-08-11-propertylab-ai-evaluation-design.md`  
**Fresh-agent context:** `docs/superpowers/handoffs/2026-08-11-propertylab-ai-evaluation-context.md`

**Target base:** `dev-chen` at `/Users/dadadineiyou/Documents/GitHub/petav3-dev-chen-integration`. Create an isolated worktree/branch named `codex/propertylab-ai-evaluation`; do not implement directly in the served `dev-chen` checkout. Do not push, merge, deploy, migrate a shared database, or mutate GPU infrastructure without a separate orchestrator decision.

---

## Frozen contracts

### C1. Scope and language

- Only PropertyLab AI output is evaluated.
- Existing four sections remain `summary`, `customer`, `sales_performance`, `meeting_report`.
- Sales-facing unsupported-content question uses `No issue`, `Issue found`, `Not sure`; never require the word “hallucination.”
- Automatic evidence verdicts use `supported`, `contradicted`, `no_evidence`.
- Automatic flags and human-confirmed issues are distinct metrics.
- V1+V2 ships the collection/evaluation/dashboard workflow; it never trains or deploys a model.

### C2. Authorization

- `/manage/zoom/evaluation` and evaluation run controls require `auth`, admin, `VIEW_ZOOM`, `super-admin` middleware, plus controller-level `isSuperAdmin()` checks.
- Existing recording visibility rules govern salesperson review and claim confirmation.
- Aggregate list responses omit transcripts, raw AI request bodies, API keys, and unrestricted reviewer notes.

### C3. Generation identity and versioning

- `zoom_meeting_analyses` remains the latest meeting/variant row.
- `zoom_meeting_analysis_generations` is append-only and represents one logical analysis job. Its UUID is created in the queued job constructor and remains stable across retries.
- `zoom_meeting_analysis_generation_attempts` records each queue attempt separately; provider-internal HTTP retries remain unavailable until transports expose them.
- A successful generation is tied to the exact `ai_request_id`, provider/model/revision/image digest, resolved prompt SHA-256, schema/retrieval/taxonomy versions, transcript and analyzed-input SHA-256, generation parameters, application commit, and normalized output snapshot.
- `conversation-analysis-v1` is the initial schema version.
- Prompt SHA is computed from the exact resolved system message stored in `ai_requests.request`, not from a mutable filename or prompt key alone.

### C4. Review identity

- Existing section review identity remains analysis + generation request + section + reviewer.
- New authoritative `unsupported_content_status=no_issue|issue|unsure` drives the UI and KPI. Existing `has_hallucination=false|true|null` is retained and dual-written for rollback compatibility.
- No row means unanswered.
- Claim confirmations are separate rows unique by claim evaluation + reviewer; one person never overwrites another person’s opinion.

### C5. Evaluator integrity

- The evaluator provider is the configured Cloud AI conversation provider/model captured when the run is created.
- Run creation fails before dispatch if the provider is `peta`.
- Automatic evaluation is default-off and must pass configured provider/region policy, per-run limits, and daily request/token/vendor-cost ceilings.
- Candidate generation IDs, evaluator provider/model, prompt SHA, and schema version are frozen per run.
- Review-unit extraction is deterministic; Cloud AI classifies supplied units and cannot invent the unit list.
- Supported/contradicted evidence must be an exact substring of the authorized source input with matching character offsets; invalid evidence is stored as `no_evidence` with a validation marker.
- Current transcript storage does not support reliable timestamp or speaker claims; do not fabricate them.

### C6. Metrics

- Human review coverage, score averages, pass rate, human-confirmed unsupported rate, automatic flag rate, first-pass JSON/schema rates, failure rate, and P95 end-to-end time use the formulas in the approved design.
- `Not sure` and unanswered are excluded from the confirmed issue denominator and shown separately.
- Missing telemetry is `null/unavailable`, never zero.
- Cloud evaluator vendor cost and PropertyLab GPU cost estimates are separate.
- Any per-meeting GPU figure derived from duration is labelled an estimate; only a verified benchmark run window may be labelled allocated actual run cost.

### C7. Guidance

- Guidance is a pure, deterministic, versioned rule set; initial rule version is `propertylab-guidance-v1`.
- Ten meetings with 50% coverage may show `pilot_signal` only. A qualified recommendation needs at least 30 completed generations and 100 human-confirmed review units.
- Prompt, RAG, and fine-tune rules exactly follow the approved design thresholds; fine-tune requires 500 adjudicated development examples plus 100 held-out examples.
- Every recommendation returns rule version, sample size, triggered metrics, threshold, explanation, and candidate case IDs.

### C8. Cost and infrastructure safety

- No browser-to-vLLM metrics call and no Laravel-to-GPU SSH.
- Metrics base URL is fixed server config and cannot be supplied in HTTP input.
- No ECS start/stop/release/resize, security-group change, purchase, or public port exposure is part of this plan.
- No secret appears in Git, logs, test fixtures, Vue props, exception text, or the handoff report.

---

## Task 0: Isolate execution and prove the baseline

**Files:** no source edits.

- [ ] From `/Users/dadadineiyou/Documents/GitHub/petav3-dev-chen-integration`, verify `dev-chen` is clean and note its HEAD.
- [ ] Check whether `.worktrees/` or `worktrees/` is already preferred and ignored. If neither exists, use an external worktree under `/Users/dadadineiyou/.codex/worktrees/`.
- [ ] Create branch `codex/propertylab-ai-evaluation` from the verified `dev-chen` HEAD in a new worktree.
- [ ] Before tests or edits, assert that every path marked **Modify** in Tasks 1–9 exists on the real `dev-chen` HEAD; print all missing paths and halt rather than improvising a replacement.
- [ ] Reuse dependencies through the repository-approved mechanism; do not copy `.env` into Git or print it.
- [ ] Copy the approved design, plan, and context handoff into the worktree and force-add them because `docs/superpowers/` is generally ignored.
- [ ] Create `.superpowers/sdd/progress.md` with task status, baseline HEAD, test commands, commits, reviews, and blockers. Keep `.superpowers/` untracked.
- [ ] Run baseline focused tests:

```bash
herd php artisan test \
  tests/Feature/Zoom/ZoomPropertyLabAnalysisTest.php \
  tests/Feature/Zoom/ZoomPropertylabReviewTest.php
npm test -- \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js \
  resources/js/Components/RecordingDetail/RecordingDetail.review.test.js \
  resources/js/Components/RecordingDetail/AnalysisProviderToggle.test.js
```

Expected: exit 0. Record known warnings separately; do not call warnings failures.

- [ ] Commit docs only:

```bash
git add -f docs/superpowers/specs/2026-08-11-propertylab-ai-evaluation-design.md \
  docs/superpowers/plans/2026-08-11-propertylab-ai-evaluation.md \
  docs/superpowers/handoffs/2026-08-11-propertylab-ai-evaluation-context.md
git commit -m "docs(ai): plan PropertyLab evaluation system"
```

**Verification:** clean tracked status after the docs commit; `.superpowers/` is the only allowed untracked path.

---

## Task 1: Replace technical hallucination input with tri-state business wording

**Files:**

- Create: `database/migrations/2026_08_11_110000_add_unsupported_status_to_zoom_analysis_reviews.php`
- Modify: `app/Http/Requests/Manage/Zoom/StorePropertylabReviewRequest.php`
- Modify: `app/Http/Controllers/Manage/Zoom/RecordingsController.php`
- Modify: `src/Zoom/ZoomMeetingAnalysisReview.php`
- Modify: `src/Zoom/Repositories/ZoomMeetingAnalysisReviewRepository.php`
- Modify: `src/Zoom/Services/ZoomRecordingDetailBuilder.php`
- Modify: `resources/js/Components/RecordingDetail/PropertylabReviewCard.vue`
- Modify: `resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js`
- Modify: `tests/Feature/Zoom/ZoomPropertylabReviewTest.php`
- Create: `tests/Feature/Database/ZoomAnalysisReviewTriStateMigrationTest.php`

### Step 1 — RED backend contract

- [ ] Add tests proving `unsupported_content_status` accepts `no_issue`, `issue`, and `unsure`, rejects omitted/invalid values, and dual-writes `has_hallucination=false|true|null` respectively.
- [ ] Add migration upgrade/rollback coverage proving old boolean rows are preserved, unambiguously backfilled to the matching new status, and the legacy column becomes nullable.
- [ ] Run:

```bash
herd php artisan test \
  tests/Feature/Zoom/ZoomPropertylabReviewTest.php \
  tests/Feature/Database/ZoomAnalysisReviewTriStateMigrationTest.php
```

Expected RED: the new status is rejected and/or silently omitted because the schema/write/read path does not yet support it.

### Step 2 — GREEN backend

- [ ] Add a forward migration with authoritative status plus a nullable legacy boolean. `down()` is deterministic and preserves every unambiguous legacy value.
- [ ] Validate the new required status allow-list; temporarily accept the legacy boolean payload for rollback compatibility. Leave stale-generation/section/reviewer constraints unchanged.
- [ ] Update explicit controller mapping, repository upsert fields, model fillable/casts/presenter, and detail-builder hydration. The new status wins when both payloads are supplied; the presenter emits both fields during the compatibility window.
- [ ] Rerun the focused backend tests; expected exit 0.

### Step 3 — RED/GREEN frontend

- [ ] Add component tests for the plain-language question, three choices, explicit untouched-vs-not-sure state, payload values, saved state, and no auto-submit.
- [ ] Run the component test and observe RED before editing the component.
- [ ] Implement explicit tri-state selection so an untouched review is unanswered while `Not sure` saves `unsupported_content_status=unsure` and legacy `has_hallucination=null`.
- [ ] Rerun component test; expected exit 0.

### Step 4 — Commit and review

```bash
git add database/migrations/2026_08_11_110000_add_unsupported_status_to_zoom_analysis_reviews.php \
  app/Http/Requests/Manage/Zoom/StorePropertylabReviewRequest.php \
  app/Http/Controllers/Manage/Zoom/RecordingsController.php \
  src/Zoom/ZoomMeetingAnalysisReview.php \
  src/Zoom/Repositories/ZoomMeetingAnalysisReviewRepository.php \
  src/Zoom/Services/ZoomRecordingDetailBuilder.php \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.vue \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js \
  tests/Feature/Zoom/ZoomPropertylabReviewTest.php \
  tests/Feature/Database/ZoomAnalysisReviewTriStateMigrationTest.php
git commit -m "feat(ai): make unsupported-content reviews tri-state"
```

- [ ] Independent reviewer checks backward compatibility, null semantics, accessibility, and payload fidelity.

---

## Task 2: Capture immutable generation lineage and first-pass reliability

**Files:**

- Create: `database/migrations/2026_08_11_120000_create_zoom_analysis_generations_table.php`
- Create: `src/Zoom/ZoomMeetingAnalysisGeneration.php`
- Create: `src/Zoom/ZoomMeetingAnalysisGenerationAttempt.php`
- Create: `src/Zoom/Repositories/ZoomMeetingAnalysisGenerationRepository.php`
- Create: `src/Ai/Services/AiExecutionFingerprintService.php`
- Create: `src/Conversation/ConversationAnalysisValidator.php`
- Modify: `src/Conversation/ConversationAnalyzer.php`
- Modify: `app/Jobs/Ai/AnalyzeZoomMeeting.php`
- Modify: `src/Zoom/ZoomMeeting.php`
- Modify: `src/Ai/AiRequest.php` only if a relationship or metadata constant is needed
- Create: `tests/Unit/Conversation/ConversationAnalysisValidatorTest.php`
- Create: `tests/Feature/Zoom/ZoomAnalysisGenerationTest.php`
- Modify: `tests/Feature/Zoom/ZoomPropertyLabAnalysisTest.php`

### Step 1 — Freeze the schema in tests

- [ ] Write validator unit tests for: valid canonical raw JSON; fenced/extracted JSON not counted as first-pass parse; missing/wrong types fail strict schema; enum/range/list violations fail; normalization remains available but does not turn a failed first pass into success.
- [ ] Add a schema-drift test binding the canonical field contract to `ConversationAnalysis::SCHEMA_VERSION`; a field-contract change without a version change must fail.
- [ ] Write feature tests proving a successful PropertyLab job creates one immutable generation with exact meeting/variant/request/provider/model/revision/image digest, canonical execution fingerprint, prompt SHA, schema/retrieval/taxonomy versions, transcript/input hashes, normalized snapshot, parse/schema flags, timings, tokens, and cost.
- [ ] Test forced re-analysis produces a second generation while the latest projection points to the newest request.
- [ ] Test idempotent settlement does not duplicate a generation for the same `ai_request_id`.
- [ ] Test one stable generation UUID survives transient retries, each queue attempt is recorded, and a terminal failure settles that generation without pretending schema success.
- [ ] Test early no-op exits (meeting missing, already done without force, empty transcript) create no generation or attempt rows.
- [ ] Run the named tests; expected RED because the model/table/validator do not exist.

### Step 2 — Implement minimally

- [ ] Add named-class generation and generation-attempt migrations with no schema-level foreign keys, indexed meeting/variant/provider/model/status/completed time, unique nullable final `ai_request_id`, and unique generation+attempt number. Add a nullable current-generation pointer to the latest projection and update it transactionally so an older retry cannot replace a newer success.
- [ ] Persist both `meeting_transcript_sha256` and `analysis_input_sha256` because current analysis may prepend prior-meeting context. Persist hashes/byte lengths only; do not duplicate transcript text.
- [ ] Add strict, pure `ConversationAnalysisValidator::validateRaw(array): array` returning `{valid, errors}` and a public schema version constant.
- [ ] Extend `ConversationAnalyzer::analyzeWithMetadata()` to expose raw response text/decoded JSON and strict first-pass flags without changing `analyze()` or normalized output contracts.
- [ ] Implement canonical `AiExecutionFingerprintService` over all frozen lineage fields. Record the exact resolved system prompt SHA by loading the settled `ai_requests.request` row and hashing its system message. If unavailable, store null and expose a missing-lineage flag; never guess a prompt/model revision.
- [ ] Capture queued time in the job constructor, started time immediately before the provider call, completed time after settlement, provider duration from `ai_requests`, and end-to-end duration from queued/completed timestamps.
- [ ] Add the generation UUID as a promoted `AnalyzeZoomMeeting` constructor property created before dispatch. Never derive it inside `run()` or from the queue job UUID; `failed()` must settle that exact UUID.
- [ ] Store a logical failed generation at terminal failure; keep per-HTTP-attempt details in existing `ai_requests`.
- [ ] Rerun focused tests; expected exit 0.

### Step 3 — Regression and commit

```bash
herd php artisan test \
  tests/Unit/Conversation/ConversationAnalysisValidatorTest.php \
  tests/Feature/Zoom/ZoomAnalysisGenerationTest.php \
  tests/Feature/Zoom/ZoomPropertyLabAnalysisTest.php
```

```bash
git add database/migrations/2026_08_11_120000_create_zoom_analysis_generations_table.php \
  src/Zoom/ZoomMeetingAnalysisGeneration.php src/Zoom/ZoomMeetingAnalysisGenerationAttempt.php \
  src/Zoom/Repositories/ZoomMeetingAnalysisGenerationRepository.php \
  src/Ai/Services/AiExecutionFingerprintService.php \
  src/Conversation/ConversationAnalysisValidator.php \
  src/Conversation/ConversationAnalyzer.php app/Jobs/Ai/AnalyzeZoomMeeting.php \
  src/Zoom/ZoomMeeting.php src/Ai/AiRequest.php \
  tests/Unit/Conversation/ConversationAnalysisValidatorTest.php \
  tests/Feature/Zoom/ZoomAnalysisGenerationTest.php \
  tests/Feature/Zoom/ZoomPropertyLabAnalysisTest.php
git commit -m "feat(ai): record immutable analysis generations"
```

- [ ] Independent reviewer verifies exact prompt hashing, raw-vs-normalized semantics, retry/idempotency behavior, and no transcript duplication.

---

## Task 3: Persist deterministic review units, evaluator runs, and per-user claim opinions

**Files:**

- Create: `database/migrations/2026_08_11_130000_create_zoom_analysis_evaluation_tables.php`
- Create: `src/Zoom/ZoomAnalysisEvaluationRun.php`
- Create: `src/Zoom/ZoomAnalysisClaimEvaluation.php`
- Create: `src/Zoom/ZoomAnalysisClaimReview.php`
- Create: `src/Zoom/Services/PropertyLabReviewUnitExtractor.php`
- Create: `src/Zoom/Repositories/ZoomAnalysisEvaluationRepository.php`
- Create: `tests/Unit/Zoom/PropertyLabReviewUnitExtractorTest.php`
- Create: `tests/Feature/Zoom/ZoomAnalysisEvaluationPersistenceTest.php`

### Step 1 — RED contracts

- [ ] Unit-test the frozen extraction map:
  - `summary` scalar;
  - `customer.needs[]`, `customer.concerns[]`, and evidence-bearing customer scalar facts;
  - `sales_performance.strengths[]`, `sales_performance.improvements[]` but not coaching score;
  - `meeting_report.summary`, `key_points[]`, `next_steps[]`, and action bodies/reasons;
  - empty/non-string values omitted;
  - stable path, ordinal, section, text, and key for identical generation input.
- [ ] Feature-test run fields/counters/fingerprints, evaluator-attempt rows, unique run+unit keys, bounded excerpts/reasons, evaluator request linkage, and multiple reviewers writing separate claim-review rows.
- [ ] Test DB constraints reject duplicate reviewer opinions without overwriting another reviewer.
- [ ] Test human issue category is accepted only with `verdict=problem`, uses the frozen taxonomy, and never copies the evaluator category implicitly.
- [ ] Test meeting-grouped evaluation-cohort assignment prevents any generation from the same meeting entering both development and holdout cohorts.
- [ ] Run tests; expected RED.

### Step 2 — GREEN persistence

- [ ] Add a named-class migration for run ledger, evaluator-attempt, claim evaluation, and claim review tables with short explicit index names and no schema-level foreign keys, matching the adjacent review migration. Add `evaluation_cohort=development|holdout` to generations and nullable `human_issue_category` to claim reviews.
- [ ] Store human verdict only in `zoom_analysis_claim_reviews`; never mutate evaluator verdict/evidence to reflect a person’s opinion.
- [ ] Implement extractor as a pure deterministic service with bounded text normalization and stable SHA-256 keys.
- [ ] Implement repository methods for creating frozen runs, idempotently replacing/settling only the current run’s units, atomic counters, and per-user review upserts.
- [ ] Rerun tests; expected exit 0.

### Step 3 — Commit and review

```bash
git add database/migrations/2026_08_11_130000_create_zoom_analysis_evaluation_tables.php \
  src/Zoom/ZoomAnalysisEvaluationRun.php src/Zoom/ZoomAnalysisClaimEvaluation.php \
  src/Zoom/ZoomAnalysisClaimReview.php \
  src/Zoom/Services/PropertyLabReviewUnitExtractor.php \
  src/Zoom/Repositories/ZoomAnalysisEvaluationRepository.php \
  tests/Unit/Zoom/PropertyLabReviewUnitExtractorTest.php \
  tests/Feature/Zoom/ZoomAnalysisEvaluationPersistenceTest.php
git commit -m "feat(ai): persist evidence evaluation runs"
```

- [ ] Independent reviewer checks data invariants, history preservation, PII bounds, and per-user ownership.

---

## Task 4: Implement the independent Cloud AI evidence evaluator and capped run controls

**Files:**

- Create: `resources/prompts/propertylab_evidence_evaluation.md`
- Modify: `config/ai_prompts.php`
- Modify: `src/Ai/AiRequest.php`
- Create: `src/Zoom/Services/PropertyLabEvidenceEvaluation.php`
- Create: `src/Zoom/Services/TranscriptEvidenceRetriever.php`
- Create: `src/Zoom/PropertyLabEvaluationTaxonomy.php`
- Create: `app/Jobs/Ai/EvaluatePropertyLabGeneration.php`
- Create: `app/Http/Requests/Manage/Zoom/EvaluationRunPreviewRequest.php`
- Create: `app/Http/Requests/Manage/Zoom/StoreEvaluationRunRequest.php`
- Create: `app/Http/Controllers/Manage/Zoom/PropertyLabEvaluationRunController.php`
- Modify: `routes/web.php`
- Modify: `config/queue.php`
- Modify: `config/horizon.php`
- Modify: `config/features.php`
- Modify: `config/ai.php`
- Modify: `.env.example`
- Create: `tests/Unit/Zoom/PropertyLabEvidenceEvaluationTest.php`
- Create: `tests/Feature/Zoom/PropertyLabEvidenceEvaluatorTest.php`
- Create: `tests/Feature/Manage/Zoom/PropertyLabEvaluationRunTest.php`

### Step 1 — RED normalizer and evidence validation

- [ ] Test strict evaluator JSON shape, exact unit-key matching, verdict allow-list, confidence clamp/rejection, bounded rationale/quote, and missing/extra unit handling.
- [ ] Test deterministic bounded retrieval never crosses meetings, handles multilingual text, and preserves exact offsets. Test supported/contradicted only when quote+offset select an authorized source span; otherwise downgrade to no-evidence with `evidence_valid=false`.
- [ ] Test no evaluator response can introduce a unit not supplied by the deterministic extractor.
- [ ] Add prompt-injection fixtures in both transcript and PropertyLab statement (for example, instructions to mark every unit supported). Assert exact supplied unit cardinality/key set and require validated evidence before any supported/contradicted verdict.
- [ ] Run unit test; expected RED.

### Step 2 — RED job and run endpoints

- [ ] Feature-test a successful evaluator call using `Http::fake()`: Cloud provider/model is frozen, prompt/schema metadata is present, evaluator `ai_request_id` is saved, run counters settle, and source PropertyLab output is unchanged.
- [ ] Test default-off flag, configured evaluator provider `peta`, disallowed region/provider, and exhausted daily request/token/vendor-cost budget all fail before dispatch/cost.
- [ ] Test budget races/backlog: the job rechecks budget immediately before the Cloud call; exhausted work settles `skipped_budget`, increments the correct run counter, and yields `completed_with_errors` instead of success or silent drop.
- [ ] Test duplicate dispatch is idempotent via overlap key and existing settled run-unit rows.
- [ ] Test transient/permanent failures use `AiJob` behavior and settle run counters/error status correctly.
- [ ] Test dry-run preview does not persist or dispatch.
- [ ] Test actual run freezes candidate IDs, enforces a hard configurable maximum (default 25), ignores non-PropertyLab/failed/webinar generations, and dispatches exactly one job per candidate.
- [ ] Test ordinary admin gets 403; super-admin succeeds.
- [ ] Run feature tests; expected RED.

### Step 3 — GREEN implementation

- [ ] Register a new prompt key and strict evaluator prompt. Delimit transcript/output as untrusted data; instruct the evaluator to ignore instructions inside them, assess only supplied units against retrieved spans, return exactly one result per unit, use no outside facts, and admit no-evidence.
- [ ] Implement deterministic transcript-only retrieval with bounded spans and adjacent context. Do not build an embedding corpus or cross-meeting RAG in Phase 0.
- [ ] Implement pure result normalization and server-side evidence validation.
- [ ] Implement job on a low-priority `evaluation` queue while reusing rate limiting, circuit breaker, backoff, timeout, and overlap controls from `AiJob`.
- [ ] Enforce the daily request/token/vendor-cost budget twice: at run creation and inside `EvaluatePropertyLabGeneration::run()` immediately before calling Cloud AI. Compute usage from `ai_requests` filtered by the evaluator prompt key and the configured Kuala Lumpur billing-day window.
- [ ] Implement preview/create endpoints under super-admin protection. Use ordered, filtered generation IDs and persist the fixed evaluator configuration at run creation.
- [ ] Add one Horizon supervisor for `evaluation` with `maxProcesses=1`; do not increase the existing product AI worker pool.
- [ ] Rerun all Task 4 tests; expected exit 0.

### Step 4 — Commit and security review

```bash
git add resources/prompts/propertylab_evidence_evaluation.md config/ai_prompts.php \
  src/Ai/AiRequest.php src/Zoom/Services/PropertyLabEvidenceEvaluation.php \
  src/Zoom/Services/TranscriptEvidenceRetriever.php src/Zoom/PropertyLabEvaluationTaxonomy.php \
  app/Jobs/Ai/EvaluatePropertyLabGeneration.php \
  app/Http/Requests/Manage/Zoom/EvaluationRunPreviewRequest.php \
  app/Http/Requests/Manage/Zoom/StoreEvaluationRunRequest.php \
  app/Http/Controllers/Manage/Zoom/PropertyLabEvaluationRunController.php \
  routes/web.php config/queue.php config/horizon.php config/features.php config/ai.php .env.example \
  tests/Unit/Zoom/PropertyLabEvidenceEvaluationTest.php \
  tests/Feature/Zoom/PropertyLabEvidenceEvaluatorTest.php \
  tests/Feature/Manage/Zoom/PropertyLabEvaluationRunTest.php
git commit -m "feat(ai): evaluate PropertyLab evidence with Cloud AI"
```

- [ ] Independent security reviewer checks secret flow, provider independence, transcript scope, prompt injection resistance, SSRF boundaries, queue/cost caps, and authorization.

---

## Task 5: Show evidence suggestions and collect every reviewer’s opinion

**Files:**

- Create: `app/Http/Requests/Manage/Zoom/StorePropertyLabClaimReviewRequest.php`
- Create: `app/Http/Controllers/Manage/Zoom/PropertyLabClaimReviewController.php`
- Modify: `routes/web.php`
- Modify: `src/Zoom/Services/ZoomRecordingDetailBuilder.php`
- Modify: `resources/js/Components/RecordingDetail/PropertylabReviewCard.vue`
- Modify: `resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js`
- Modify: `resources/js/Components/RecordingDetail/RecordingDetail.review.test.js`
- Create: `tests/Feature/Manage/Zoom/PropertyLabClaimReviewTest.php`

### Step 1 — RED authorization and payload tests

- [ ] Test detail props return only claim evaluations for the current PropertyLab generation and current section, plus only the viewer’s own claim-review selection.
- [ ] Test another reviewer’s choice is not exposed or overwritten.
- [ ] Test a visible-meeting user can save `supported|problem|unsure`; invalid verdicts fail validation.
- [ ] Test stale generation, wrong section, non-visible meeting, and non-PropertyLab claim return 403/409/422 according to existing controller conventions.
- [ ] Test saving a claim opinion never calls AI, dispatches a job, changes analysis status, or sends notifications.
- [ ] Run feature test; expected RED.

### Step 2 — RED/GREEN Vue behavior

- [ ] Add component tests for evaluator loading/empty/failed states, suggested verdict/reason/evidence, plain-language confirmation choices, explicit Save, busy/double-submit protection, and stale refresh callback.
- [ ] Verify RED before implementation.
- [ ] Render suggestions below the matching section review; do not render full transcript or unrelated sections.
- [ ] Preserve existing review scores and save flow unchanged.
- [ ] Rerun frontend tests; expected exit 0.

### Step 3 — GREEN backend and commit

- [ ] Implement Form Request/controller/repository call with existing Zoom visibility and current-generation checks.
- [ ] Extend detail builder without widening other reviewers’ data.
- [ ] Rerun backend and frontend focused tests.

```bash
git add app/Http/Requests/Manage/Zoom/StorePropertyLabClaimReviewRequest.php \
  app/Http/Controllers/Manage/Zoom/PropertyLabClaimReviewController.php routes/web.php \
  src/Zoom/Services/ZoomRecordingDetailBuilder.php \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.vue \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js \
  resources/js/Components/RecordingDetail/RecordingDetail.review.test.js \
  tests/Feature/Manage/Zoom/PropertyLabClaimReviewTest.php
git commit -m "feat(ai): collect evidence review opinions"
```

- [ ] Independent reviewer checks data isolation, stale-generation behavior, accessibility, and that automatic/human verdicts remain visually distinct.

---

## Task 6: Build deterministic metrics and training guidance

**Files:**

- Create: `src/Zoom/Services/PropertyLabEvaluationMetrics.php`
- Create: `src/Zoom/Services/PropertyLabTrainingGuidance.php`
- Create: `database/migrations/2026_08_11_135000_create_propertylab_training_guidance_snapshots_table.php`
- Create: `src/Zoom/PropertyLabTrainingGuidanceSnapshot.php`
- Create: `tests/Unit/Zoom/PropertyLabEvaluationMetricsTest.php`
- Create: `tests/Unit/Zoom/PropertyLabTrainingGuidanceTest.php`

### Step 1 — RED formula tests

- [ ] Use fixed fixtures to test every formula and denominator in design section 6, including zero-data/null values, not-sure exclusion, nearest-rank percentile boundaries, multiple reviewers, multiple generations, and automatic-vs-human separation. Automatic flag rate includes transcript-verifiable units only; judgment/advice and external-fact units are separate counts.
- [ ] Test cost fields: evaluator vendor cost; duration-based PropertyLab GPU estimate; optional verified benchmark allocation; missing price/window returns null.
- [ ] Run unit tests; expected RED.

### Step 2 — RED guidance table tests

- [ ] Test `collect_more_evidence` below 10 meetings/50% coverage, `pilot_signal` at the pilot threshold, and no qualified recommendation below 30 completed generations plus 100 human-confirmed units.
- [ ] Test prompt recommendation for schema/instruction/format/transcript-evidence patterns.
- [ ] Test RAG requires at least 30 human-confirmed external-knowledge errors, at least 60% of actionable errors, and a declared permission-filtered authoritative corpus.
- [ ] Test fine-tuning requires 500 adjudicated development examples, 100 held-out examples, error ≥10% across two prompt versions, schema ≥99%, and prompt/RAG prerequisites; holdout rows are never training candidates.
- [ ] Test mixed signals return multiple ranked candidates without claiming certainty.
- [ ] Test every result contains rule version, numerator/denominator, sample, threshold, explanation, and case IDs.
- [ ] Test distinct prompt SHA-256 values are ordered by first-seen generation time when numeric prompt versions are unavailable.
- [ ] Run unit tests; expected RED.

### Step 3 — GREEN pure services

- [ ] Implement DB-agnostic calculators over typed arrays/collections. Keep query construction out of the rule engine so formulas are directly testable.
- [ ] Use the explicit nearest-rank sorting contract for P95 and mark it unstable below 20 samples; do not interpolate.
- [ ] Implement guidance rules exactly; do not add ML-based recommendations or hidden thresholds. Persist explicit super-admin-created snapshots with filters, metrics, ruleset, reasons, and candidate IDs.
- [ ] Rerun tests; expected exit 0.

### Step 4 — Commit and domain review

```bash
git add database/migrations/2026_08_11_135000_create_propertylab_training_guidance_snapshots_table.php \
  src/Zoom/PropertyLabTrainingGuidanceSnapshot.php \
  src/Zoom/Services/PropertyLabEvaluationMetrics.php \
  src/Zoom/Services/PropertyLabTrainingGuidance.php \
  tests/Unit/Zoom/PropertyLabEvaluationMetricsTest.php \
  tests/Unit/Zoom/PropertyLabTrainingGuidanceTest.php
git commit -m "feat(ai): derive PropertyLab training guidance"
```

- [ ] Independent reviewer recomputes fixture outputs by hand and checks that a 10-meeting pilot cannot trigger fine-tuning.

---

## Task 7: Add safe GPU telemetry collection and cost inputs

**Files:**

- Create: `database/migrations/2026_08_11_140000_create_peta_gpu_metric_samples_table.php`
- Create: `src/Ai/PetaGpuMetricSample.php`
- Create: `src/Ai/Services/PetaGpuMetricsCollector.php`
- Create: `app/Jobs/Ai/CollectPetaGpuMetrics.php`
- Modify: `config/ai.php`
- Modify: `.env.example`
- Modify: `app/Console/Kernel.php`
- Create: `tests/Unit/Ai/PetaGpuMetricsCollectorTest.php`
- Create: `tests/Feature/Ai/CollectPetaGpuMetricsTest.php`

### Step 1 — RED collector safety tests

- [ ] Test blank configuration performs no network call and records/returns unavailable without throwing.
- [ ] Test URL comes only from config, allows HTTPS or explicitly configured private HTTP hosts, rejects redirects/public-host drift, uses short timeout, and sends a redacted read-only credential.
- [ ] Test content-type/size limits, malformed payload, partial metrics, counter parsing, and stale timestamps.
- [ ] Test allow-list persistence excludes unknown/high-cardinality labels and never stores raw payload, prompt, transcript, customer, or request IDs.
- [ ] Run tests; expected RED.

### Step 2 — GREEN collector/job/schedule

- [ ] Add a named-class migration with no schema-level foreign keys, nullable metric columns, and collection status/error fields. Absence remains null.
- [ ] Add config keys for enabled flag, fixed base URL, read-only token, timeout, maximum bytes, stale-after seconds, GPU hourly USD, and optional fixed hourly USD. All default disabled/null except the documented Phase 0 reference price may appear only as an example value.
- [ ] Implement a fail-soft collector and queued job. Logs contain status/error class but never URL credentials or raw payload.
- [ ] Schedule collection every minute with `withoutOverlapping()` only when enabled; no automatic infrastructure change.
- [ ] Keep collection disabled unless a documented private Laravel-to-metrics route exists. No private route is assumed in Phase 0; a permanently unavailable/default-off state is an acceptable verified result and must not trigger per-minute network traffic.
- [ ] Rerun tests; expected exit 0.

### Step 3 — Commit and security review

```bash
git add database/migrations/2026_08_11_140000_create_peta_gpu_metric_samples_table.php \
  src/Ai/PetaGpuMetricSample.php src/Ai/Services/PetaGpuMetricsCollector.php \
  app/Jobs/Ai/CollectPetaGpuMetrics.php config/ai.php .env.example \
  app/Console/Kernel.php tests/Unit/Ai/PetaGpuMetricsCollectorTest.php \
  tests/Feature/Ai/CollectPetaGpuMetricsTest.php
git commit -m "feat(ai): collect private GPU telemetry"
```

- [ ] Independent security reviewer checks SSRF, secret redaction, response bounds, scheduling, missing-data semantics, and confirms no public port or SSH path was introduced.

---

## Task 8: Build the super-admin Evaluation dashboard and run UI

**Files:**

- Create: `app/Http/Requests/Manage/Zoom/PropertyLabEvaluationQueryRequest.php`
- Create: `app/Http/Controllers/Manage/Zoom/PropertyLabEvaluationController.php`
- Create: `src/Zoom/Repositories/PropertyLabEvaluationReadRepository.php`
- Modify: `routes/web.php`
- Modify: `resources/js/Components/SectionTabs.vue`
- Create: `resources/js/Pages/Manage/Zoom/Evaluation/Index.vue`
- Create: `resources/js/Pages/Manage/Zoom/Evaluation/Index.test.js`
- Create: `tests/Feature/Manage/Zoom/PropertyLabEvaluationDashboardTest.php`

### Step 1 — RED backend authorization/query tests

- [ ] Test super-admin renders the new Inertia component and ordinary admin receives 403.
- [ ] Test period whitelist 7/30/90, provider/model/prompt/schema/status/section/review-completeness filters, pagination, and empty state.
- [ ] Test KPI/trend/guidance props match fixed database fixtures and do not include transcripts, raw request/response bodies, keys, or unrestricted notes.
- [ ] Test automatic and human-confirmed rates are separate and missing GPU data is unavailable/stale.
- [ ] Test case links use the existing recording detail `?detail={uuid}` path.
- [ ] Run feature test; expected RED.

### Step 2 — GREEN backend

- [ ] Add `GET /manage/zoom/evaluation` inside the Zoom group with explicit `super-admin` middleware and controller guard.
- [ ] Implement one scoped read repository with bounded selects/eager loads; no N+1 queries and no full transcript fetch for aggregate/list requests.
- [ ] Build period trends with zero-filled buckets following `ZoomDashboardController` conventions.
- [ ] Feed typed aggregates into the Task 6 metrics/guidance services.
- [ ] Rerun backend test; expected exit 0.

### Step 3 — RED/GREEN frontend

- [ ] Add `Evaluation` Zoom section tab with `superAdmin: true`; do not add a left-sidebar item.
- [ ] Test period/filter query synchronization, KPI labels/denominators, guidance evidence, evaluation-run preview/confirm flow, loading/empty/error states, case deep links, and stale GPU display.
- [ ] Implement a focused page using existing dashboard cards, filters, tables, and Inertia router conventions.
- [ ] Label Cloud evaluator cost, PropertyLab GPU estimate, and allocated benchmark cost distinctly.
- [ ] Test that no combined “total AI cost” is rendered and that every duration-derived PropertyLab cost carries a visible `Estimate` label in the DOM, not only a tooltip.
- [ ] Rerun frontend test; expected exit 0.

### Step 4 — Commit and review

```bash
git add app/Http/Requests/Manage/Zoom/PropertyLabEvaluationQueryRequest.php \
  app/Http/Controllers/Manage/Zoom/PropertyLabEvaluationController.php \
  src/Zoom/Repositories/PropertyLabEvaluationReadRepository.php routes/web.php \
  resources/js/Components/SectionTabs.vue \
  resources/js/Pages/Manage/Zoom/Evaluation/Index.vue \
  resources/js/Pages/Manage/Zoom/Evaluation/Index.test.js \
  tests/Feature/Manage/Zoom/PropertyLabEvaluationDashboardTest.php
git commit -m "feat(ai): add PropertyLab evaluation dashboard"
```

- [ ] Independent backend and frontend reviewers check RBAC, query scale, formula presentation, accessibility, and sensitive-data minimization.

---

## Task 9: Backfill lineage, document operations, and verify end to end

**Files:**

- Create: `app/Console/Commands/BackfillZoomAnalysisGenerations.php`
- Create: `tests/Feature/Console/BackfillZoomAnalysisGenerationsTest.php`
- Modify: `docs/operations/phase0-cloud-gpu-poc-runbook.md`
- Modify: `docs/ai-poc-evaluation-checklist.zh.md`
- Modify: `docs/ai-poc-evaluation-checklist.en.md`
- Create: `docs/operations/propertylab-ai-evaluation-runbook.md`

### Step 1 — Safe backfill TDD

- [ ] Test `--dry-run`, required/capped `--limit`, idempotency, PropertyLab-only filtering, current-generation snapshot creation from `zoom_meeting_analyses + ai_requests`, explicit lineage-missing flags, and `lineage_source=backfill` for data that cannot prove first-pass parse/schema history.
- [ ] Test the command never calls either AI provider and never dispatches evaluation unless a separate run is explicitly created.
- [ ] Implement chunked command with progress counts and no recurring schedule.
- [ ] Run command test; expected exit 0.

### Step 2 — Runbook

- [ ] Document migrations, queue worker, dashboard access, dry-run/create evaluation run, retry/failure recovery, metrics configuration, stale telemetry, cost definitions, prompt/schema/guidance version changes, privacy, export boundaries, and rollback.
- [ ] Update Phase 0 checklist with exact metric formulas and verified/unverified evidence fields.
- [ ] State that 10 meetings validate collection UX only; Prompt/RAG/fine-tune recommendations need their frozen minimum evidence.
- [ ] State GPU infrastructure changes and ECS lifecycle actions remain manual/confirmed operations.

### Step 3 — Focused and full verification

Run focused backend suite:

```bash
herd php artisan test \
  tests/Unit/Conversation/ConversationAnalysisValidatorTest.php \
  tests/Unit/Zoom/PropertyLabReviewUnitExtractorTest.php \
  tests/Unit/Zoom/PropertyLabEvidenceEvaluationTest.php \
  tests/Unit/Zoom/PropertyLabEvaluationMetricsTest.php \
  tests/Unit/Zoom/PropertyLabTrainingGuidanceTest.php \
  tests/Unit/Ai/PetaGpuMetricsCollectorTest.php \
  tests/Feature/Zoom/ZoomPropertyLabAnalysisTest.php \
  tests/Feature/Zoom/ZoomAnalysisGenerationTest.php \
  tests/Feature/Zoom/ZoomAnalysisEvaluationPersistenceTest.php \
  tests/Feature/Zoom/PropertyLabEvidenceEvaluatorTest.php \
  tests/Feature/Zoom/ZoomPropertylabReviewTest.php \
  tests/Feature/Manage/Zoom/PropertyLabClaimReviewTest.php \
  tests/Feature/Manage/Zoom/PropertyLabEvaluationRunTest.php \
  tests/Feature/Manage/Zoom/PropertyLabEvaluationDashboardTest.php \
  tests/Feature/Ai/CollectPetaGpuMetricsTest.php \
  tests/Feature/Console/BackfillZoomAnalysisGenerationsTest.php
```

Run focused frontend suite:

```bash
npm test -- \
  resources/js/Components/RecordingDetail/PropertylabReviewCard.test.js \
  resources/js/Components/RecordingDetail/RecordingDetail.review.test.js \
  resources/js/Components/RecordingDetail/AnalysisProviderToggle.test.js \
  resources/js/Pages/Manage/Zoom/Evaluation/Index.test.js
```

Run production builds:

```bash
npm run build
```

Run formatting/syntax/diff checks only on touched files using the project’s installed tools, then:

```bash
git diff --check
git status --short
```

- [ ] Run `tests/Unit/TestDatabaseGuardTest.php` before any database-mutating test. Verify all new migrations up/down on the repository-sanctioned isolated `petav3_testing` connection (or a uniquely named temporary MySQL database configured only for this run). Do not use the shared development database. If the full historical SQLite chain fails on a known unrelated MySQL-only migration, run and report the exact target migrations on an isolated MySQL-compatible test DB; never hide the unrelated failure.
- [ ] Start the isolated worktree application on a dedicated localhost port or otherwise prove the browser is serving the exact branch/HEAD before acceptance.
- [ ] Browser acceptance: super-admin access; ordinary-admin denial; tri-state review; automatic evidence display; per-user opinion isolation; dashboard filters/KPIs/guidance; dry-run without dispatch; actual run confirmation; stale GPU state; no sensitive transcript on aggregate page.
- [ ] Do not invoke a real paid Cloud evaluator during automated tests. A real 10-meeting pilot run requires the user’s separately confirmed cost window after code acceptance.

### Step 4 — Commit documentation

```bash
git add app/Console/Commands/BackfillZoomAnalysisGenerations.php \
  tests/Feature/Console/BackfillZoomAnalysisGenerationsTest.php \
  docs/operations/phase0-cloud-gpu-poc-runbook.md \
  docs/ai-poc-evaluation-checklist.zh.md docs/ai-poc-evaluation-checklist.en.md \
  docs/operations/propertylab-ai-evaluation-runbook.md
git commit -m "docs(ai): operationalize PropertyLab evaluation"
```

---

## Final review gates

- [ ] Each implementation task has a fresh worker report and independent task review in `.superpowers/sdd/`.
- [ ] A whole-branch code reviewer checks correctness, maintainability, migrations, and silent-failure behavior.
- [ ] A security reviewer checks authorization, PII, secrets, SSRF, prompt injection, queue/cost abuse, and data export.
- [ ] A database reviewer checks indexes, uniqueness, migration rollback, aggregation query plans, and append-only invariants.
- [ ] A test reviewer checks behavioral coverage, especially stale generations, per-user isolation, first-pass schema truthfulness, and formula denominators.
- [ ] Search for placeholders/TODOs and resolve any introduced by this branch.
- [ ] Verify every design acceptance criterion maps to a passing test or a documented, externally blocked real-infrastructure check.
- [ ] Claude Code receives the context handoff plus approved plan for an independent plan review before Task 0 execution, and later receives the exact approved task brief for implementation. Claude’s output is advisory until Codex verifies the diff and tests in the assigned worktree.
- [ ] Final handoff clearly separates verified results, estimates, assumptions, and unverified production/GPU claims.

