# Claude Context Handoff — PropertyLab AI Evaluation and Training Guidance

**Date:** 2026-08-11  
**Target repository:** `/Users/dadadineiyou/Documents/GitHub/petav3-dev-chen-integration`  
**Target base branch:** `dev-chen`  
**Open PR:** #110 (`dev-chen` → `master`)  
**Design source of truth:** `docs/superpowers/specs/2026-08-11-propertylab-ai-evaluation-design.md`  
**Execution plan:** `docs/superpowers/plans/2026-08-11-propertylab-ai-evaluation.md`

## Why this exists

This is a self-contained context package for Claude Code or another fresh implementation/review agent. Read it, the approved design, the execution plan, repository `AGENTS.md`, and every file named in the relevant plan task before editing.

## Product context

petaV3 is a Laravel + Inertia + Vue application used by PropertyLab. The current project compares two AI analyses of the same Zoom sales meeting:

- **Cloud AI:** the configurable external provider; the UI intentionally uses this provider-neutral label instead of “OpenAI.”
- **PropertyLab AI:** the Phase 0 private inference provider named `peta` in configuration.

The comparison must stay fair: both sides receive the identical transcript, analysis prompt, and JSON schema. Only provider/model changes.

The business wants evidence to decide whether private AI is good and economical enough to continue, and which quality problems require prompt changes, retrieval-augmented generation, or eventual fine-tuning.

## Recent completed feature

The `dev-chen` branch already contains the new comparison and personal review work. Relevant commits include:

- `3b56d7cf feat(ai): label external provider as Cloud AI`
- `1a8f486f feat(ai): review PropertyLab output by section`
- `20c6fcc5 fix(ai): guard stale review saves`
- `2e83b8f8 feat(ai): add PropertyLab review card`
- earlier commits on the same branch expose/save per-user PropertyLab reviews and create the review table.

Current behaviour:

- Four tabs: `summary`, `customer`, `sales_performance`, `meeting_report`.
- Each tab switches between Cloud AI and PropertyLab AI.
- Only PropertyLab output is reviewed.
- Each authenticated user reviews their own view of the exact PropertyLab generation.
- Review identity is `zoom_meeting_analysis_id + ai_request_id + section + reviewer_id`.
- Save is explicit. It never triggers AI, queues, regeneration, or notification.
- Stale `ai_request_id` saves are rejected.
- Webinar meetings are excluded.

## Phase 0 infrastructure facts

- Alibaba Cloud ECS instance: `i-8psbjd6ru30b79uxd2xb`
- Region: Kuala Lumpur / `ap-southeast-3`
- Shape: `ecs.gn8is.2xlarge`, 8 vCPU, 64 GiB RAM, 1× NVIDIA L20 48GB
- Ubuntu 22.04; NVIDIA driver 580.126.09; CUDA 12.8; Docker
- vLLM model name: `peta-qwen3.6-27b-fp8`
- Billing: pay-as-you-go, reference GPU instance price USD 2.641736/hour plus disk/network
- eRDMA off; one 100 GiB PL0 system disk; public IPv4 5 Mbps

Do not purchase, resize, stop, release, delete, or change fee-bearing infrastructure. Do not expose port 8000. Do not request or print passwords, API keys, private keys, or tokens. Infrastructure mutations require separate explicit user confirmation.

Existing Phase 0 evidence:

- 4 concurrent synthetic agency requests: 1,256 output tokens in 21.521s, 58.36 aggregate output tokens/s.
- Short smoke response `PETA_OK`: 0.258s.
- Real portal stream: 1,736 input / 1,127 output tokens, 64.7s, about 17.4 output tokens/s.
- Long business-shaped responses exposed output-ceiling and grounding risks.
- Current decision: conditional Go for continued Phase 0; No-Go for Phase 1/production or autonomous customer advice.

Read:

- `docs/operations/phase0-cloud-gpu-poc-runbook.md`
- `docs/ai-poc-evaluation-checklist.zh.md`
- `docs/ai-poc-evaluation-checklist.en.md`
- `docs/cloud-gpu-internal-ai-plan.md`
- `docs/ai-cloud-gpu-engineering-sop.zh.md`

## Code architecture facts

### Analysis generation

- `app/Jobs/Ai/AnalyzeZoomMeeting.php`
  - current/Cloud variant is generated first;
  - successful current analysis dispatches PropertyLab analysis;
  - PropertyLab forces provider `peta` and its configured default model.
- `src/Conversation/ConversationAnalyzer.php`
  - both variants use the same `conversation_analysis` prompt and `ConversationAnalysis` schema.
- `src/Zoom/Repositories/ZoomMeetingRepository.php`
  - saves latest variant projections.
- `database/migrations/2026_08_10_120000_create_zoom_meeting_analyses_table.php`
  - one latest row per meeting/variant; re-analysis overwrites it.

### Request observability

- `src/Ai/Services/AiClient.php`
- `src/Ai/Repositories/AiRequestRepository.php`
- `src/Ai/AiRequest.php`
- `database/migrations/2026_06_10_000003_create_ai_requests_table.php`

`ai_requests` already records request/response, provider, model, prompt key, subject, status, tokens, estimated vendor cost, duration, error, and metadata. It does not currently record prompt version, schema version, first-pass schema success, queue wait, or GPU metrics.

### Prompt/schema

- `resources/prompts/conversation_analysis.md`
- `src/Conversation/ConversationAnalysis.php`
- `config/ai_prompts.php`
- `ai_prompt_versions` history from `database/migrations/2026_07_11_120001_create_ai_prompt_tables.php`

The exact resolved prompt is present in the stored request. Numeric prompt history cannot always be reconstructed later, so capture a SHA-256 snapshot at generation time.

### Human reviews

- `src/Zoom/ZoomMeetingAnalysisReview.php`
- `src/Zoom/Repositories/ZoomMeetingAnalysisReviewRepository.php`
- `app/Http/Requests/Manage/Zoom/StorePropertylabReviewRequest.php`
- `app/Http/Controllers/Manage/Zoom/RecordingsController.php`
- `src/Zoom/Services/ZoomRecordingDetailBuilder.php`
- `resources/js/Components/RecordingDetail/PropertylabReviewCard.vue`

The current read path intentionally returns only the viewer’s reviews for the current generation. Dashboard aggregation needs a new protected cross-reviewer query and must not weaken the existing detail contract.

### Transcript limitation

`zoom_meetings.transcript` is plain text. `src/Zoom/Support/VttParser.php` removes cue timestamps/tags, and fallback transcription does not persist diarized raw segments. Evidence can use exact quote plus character offsets only. Do not invent timestamps or speakers.

### Queue/reliability patterns

- `app/Jobs/Ai/AiJob.php`: AI queue, rate limiting, circuit breaker, overlap lock, retries, timeout.
- `app/Console/Commands/AutoAnalyzeZoomRecordings.php`: capped/dry-run batch pattern.
- `app/Console/Commands/WatchZoomAiPipeline.php`: backlog/failure recovery pattern.
- `src/Analysis/Reference/CatalogSyncRun.php`: persistent run-ledger precedent.

### Dashboard/security patterns

- `routes/web.php` Zoom group under `/manage/zoom`.
- `resources/js/Components/SectionTabs.vue`: supports `superAdmin: true`.
- `app/Http/Middleware/EnsureUserIsSuperAdmin.php`.
- `app/Http/Controllers/Manage/Zoom/ZoomDashboardController.php`.
- `app/Http/Controllers/Manage/Zoom/ZoomAiAgentController.php`.
- `resources/js/Pages/Manage/Zoom/Dashboard/Index.vue`.

The new console belongs at `/manage/zoom/evaluation`, with both route middleware and controller authorization. Do not add a global sidebar entry.

## Frozen product decisions

1. Implement V1 and V2 in this project, but do not automate training.
2. Human review scores target at least 4/5 for accuracy, completeness, and actionability.
3. Human-confirmed unsupported-content rate target is at most 5%.
4. First-pass JSON schema success target is at least 99%.
5. P95 full analysis time target is at most 60 seconds.
6. Record model, prompt, schema, transcript, and evaluation lineage.
7. Use Cloud AI as the independent evidence evaluator; fail closed if configured evaluator provider is `peta`.
8. Automatic evidence output is advisory; human decisions remain separate.
9. Sales wording is `No issue / Issue found / Not sure`, not hallucination terminology.
10. Ten meetings validate the workflow, but cannot justify fine-tuning.
11. Guidance is deterministic and explains thresholds/sample sizes.
12. GPU metrics use a private fixed server-side endpoint and fail safely.
13. No recurring bulk evaluation initially; super-admin starts a capped run after a dry-run preview.
14. Keep Phase 0 single-instance scope. No K8s or multi-node design.

## Working rules for Claude

- Work only in the isolated worktree/branch assigned by the orchestrator.
- Read repository `CLAUDE.md` and `GUIDELINES.md` before edits. If an `AGENTS.md` is present, read it too; its absence is not a blocker.
- Follow the execution plan in order. Do not silently change frozen contracts.
- Use TDD: add the named failing test, verify the expected failure, implement minimally, rerun focused tests.
- Commit one logical concern per task using the repository’s conventional commit style.
- Never modify `.env`, secrets, shared database content, running GPU infrastructure, or production.
- Never run destructive database or Git commands.
- Do not edit already-committed migrations; add new forward migrations.
- Preserve user/unrelated work and do not rewrite existing comparison behaviour.
- For every task, report files changed, tests run with exact results, unresolved risks, and commit SHA.
- Stop and report if a required contract cannot be implemented without broadening scope or changing security/cost assumptions.

## Definition of done

The implementation is done only after the focused backend/frontend tests, migration verification on an isolated database, production client + SSR builds, authorization review, security review, full branch review, and browser acceptance on the correct worktree all pass or have a clearly documented external blocker.

