# PropertyLab AI Evaluation and Training Guidance — Design

**Status:** Approved scope, implementation pending independent review  
**Date:** 2026-08-11  
**Product owner decisions:** V1 and V2 ship together; evaluate only PropertyLab AI; Cloud AI is the independent evidence evaluator; recommendations are advisory and never start training automatically.

## 1. Outcome

Build a super-admin quality console for the PropertyLab AI Zoom analysis pilot. It must combine:

- human reviews from salespeople;
- automatically collected model, prompt, schema, latency, token, failure, and cost data;
- claim-to-transcript evidence checks performed by Cloud AI;
- optional human confirmation of automatically flagged claims;
- GPU/runtime telemetry when the private metrics endpoint is configured; and
- deterministic guidance that distinguishes issues best addressed by prompt changes, RAG, or fine-tuning.

The feature supports Phase 0 evidence gathering on one Alibaba Cloud L20 instance. It does not provision infrastructure, expose vLLM publicly, start training, create a model registry service, or introduce Kubernetes/multi-node architecture.

## 2. Existing baseline

The current product already provides side-by-side Cloud AI and PropertyLab AI results for four sections:

1. `summary`
2. `customer`
3. `sales_performance`
4. `meeting_report`

Both variants use the same transcript, `conversation_analysis` prompt, and normalized JSON shape. The PropertyLab variant uses the configured `peta` provider and currently targets the Qwen 27B FP8 Phase 0 deployment.

Each signed-in user can submit one review per PropertyLab generation and section. Reviews contain accuracy, completeness, and actionability scores, an unsupported-content answer, and notes. A review is bound to the exact `ai_request_id`; stale generations cannot be updated.

The current `zoom_meeting_analyses` row is a latest-state projection and is overwritten on re-analysis. `ai_requests` is append-only and already records provider, model, request/response, token counts, estimated vendor cost, duration, status, error, and metadata. The new design therefore adds an immutable generation record instead of treating the latest projection as history.

## 3. User experience

### 3.1 Sales review wording

Do not show the word “hallucination” to salespeople. The existing PropertyLab review card asks:

> Does this analysis contain anything that is not supported by what was said in the meeting?

Answers:

- `No issue`
- `Issue found`
- `Not sure`

The authoritative persisted field is `unsupported_content_status` with `no_issue`, `issue`, or `unsure`. During the compatibility window, dual-write the legacy field as:

- `false` = no issue
- `true` = issue found
- `null` on a saved review row = explicitly not sure
- no review row = unanswered

The legacy database column remains `has_hallucination` for rollback compatibility, but product copy, new APIs, and KPIs use `unsupported_content_status` and “unsupported content.”

### 3.2 Evidence assistance

Below the section review, show automatically flagged review units for the same generation and section. Each item contains:

- the PropertyLab statement;
- evaluator verdict: `Supported`, `Contradicted`, or `No evidence`;
- a short transcript excerpt when one was validated;
- a short evaluator reason and confidence; and
- an optional human decision: `Supported`, `Problem`, or `Not sure`.

Automatic verdicts are suggestions. They never overwrite human reviews and are never counted as confirmed unsupported content until a person confirms a problem.

### 3.3 Super-admin console

Add a super-admin-only `Evaluation` tab under Manage → Zoom at `/manage/zoom/evaluation`.

The page contains:

- 7/30/90-day period controls;
- quality, coverage, reliability, latency, and cost KPIs;
- quality and failure trends;
- provider/model/prompt/schema filters;
- a training-guidance panel;
- a paginated case table with deep links to the existing recording detail; and
- evaluation run/backfill controls with a dry-run preview and a hard case limit.

The aggregate page does not show full transcripts. Evidence excerpts and reviewer notes appear only in a protected case detail.

## 4. Data contracts

### 4.1 Immutable analysis generations

Create `zoom_meeting_analysis_generations` with one stable row per logical analysis job. Its UUID is generated when the job is constructed and is reused across queue retries:

- identity: UUID, meeting ID, analysis variant, final `ai_request_id` (unique when present);
- model lineage: provider, model, model revision, serving image digest, prompt key, prompt version when resolvable, prompt content SHA-256, schema/retrieval/taxonomy versions, application commit, and generation parameters;
- input lineage: meeting transcript SHA-256, exact analyzed-input SHA-256, and byte lengths; never duplicate the transcript here;
- output snapshot: normalized analysis JSON for successful attempts;
- reliability: status, first-pass JSON parse success, first-pass schema success, normalization applied, error class/message;
- timing: queued, started, completed, queue duration, provider duration, end-to-end duration;
- usage: input/output tokens and provider cost copied from `ai_requests` where available;
- timestamps.

Create `zoom_meeting_analysis_generation_attempts` for each queue attempt: generation ID, attempt number, `ai_request_id`, status/error class, and start/end timestamps. This distinguishes queue retries from provider-internal HTTP retries, which remain unavailable until the transport instruments them.

The generation row is immutable after a single idempotent settlement. `zoom_meeting_analyses` remains the latest projection used by the recording UI and gains a nullable current-generation pointer. Updating that pointer uses a transaction/compare-and-swap so an older retry cannot replace a newer successful generation.

`schema_version` starts as `conversation-analysis-v1`. It changes only when the canonical output contract changes. Prompt SHA is computed from the exact resolved system prompt sent to the provider, not from a mutable config name.

### 4.2 Evaluation run ledger

Create `zoom_meeting_analysis_evaluation_runs`:

- UUID, status (`preview`, `queued`, `running`, `completed`, `completed_with_errors`, `failed`);
- fixed filters and ordered candidate generation IDs;
- evaluator provider/model/revision, prompt SHA, evaluator schema/taxonomy version, and a canonical evaluator fingerprint;
- requested limit, total/queued/succeeded/failed/skipped counters;
- started/completed timestamps, actor, error summary;
- evaluator token/cost totals;
- PropertyLab batch wall-clock cost fields when the run is explicitly designated as a benchmark run.

Candidate IDs are frozen when a real run is created. A dry-run returns counts and candidate IDs but persists no run and dispatches no jobs. Automatic evaluation is behind a default-off feature flag, a hard per-run limit, and a hard daily request/vendor-cost budget.

### 4.3 Deterministic review units and evidence evaluations

Create `zoom_meeting_analysis_claim_evaluations` with one row per deterministic output review unit:

- run ID, generation ID, section, JSON field path, item ordinal;
- stable unit key: SHA-256 of generation UUID + path + ordinal + normalized text;
- PropertyLab statement text;
- evaluator verdict: `supported`, `contradicted`, `no_evidence`;
- issue category: `grounding`, `instruction`, `external_knowledge`, `format`, or `other`;
- evidence quote plus character start/end offsets;
- rationale and bounded confidence;
- evaluator `ai_request_id`;
- evidence validation state and source hash.

Create `zoom_meeting_analysis_claim_reviews` for human opinions:

- claim evaluation, reviewer, verdict (`supported`, `problem`, `unsure`), nullable human issue category from the frozen taxonomy, bounded note, role-at-review, and timestamp;
- unique claim evaluation + reviewer;
- a reviewer updates only their own opinion and never overwrites another person or the evaluator suggestion.

Review units are extracted deterministically from canonical fields. Array fields yield one unit per item. Narrative scalar fields yield one unit for the whole field in V1/V2; the evaluator is not allowed to invent a different claim list. Each unit is typed as transcript-verifiable, external-fact, or judgment/advice. Numeric coaching scores and pure recommendations are excluded from transcript-grounding rates because they are judgments, not transcript facts.

Evidence validation is server-side:

- supported or contradicted requires a non-empty quote that is an exact transcript substring;
- offsets must select exactly the stored quote;
- invalid evidence is downgraded to `no_evidence` and recorded as a validation failure;
- transcript timestamps and speaker identity are not fabricated because current transcript persistence does not reliably retain them.

The table stores only short, bounded excerpts. Full transcript remains in the existing Zoom record and access path.

### 4.4 GPU metric samples

Create `peta_gpu_metric_samples` for allow-listed, low-cardinality snapshots:

- source/instance label from configuration;
- collected timestamp and collection status;
- GPU utilization, VRAM used/total;
- running/queued request counts;
- prompt/generation tokens per second when available;
- request latency P50/P95 when available;
- cumulative error, OOM, and restart counters when available;
- error code/message for unavailable samples.

No prompt, transcript, customer, meeting, user, or request ID may be stored as a metric label.

The endpoint URL and read-only credential come only from server configuration. The browser never calls vLLM metrics, and Laravel never SSHes to the GPU host. Missing data is displayed as unavailable/stale, not zero.

## 5. Automatic evaluator flow

```mermaid
flowchart LR
    A["PropertyLab generation settled"] --> B["Immutable generation snapshot"]
    B --> C["Deterministic review-unit extraction"]
    C --> D["Queued Cloud AI evidence evaluator"]
    D --> E["Strict JSON validation"]
    E --> F["Server-side transcript quote validation"]
    F --> G["Suggested verdicts and evidence"]
    G --> H["Sales/admin confirmation"]
    H --> I["Quality dashboard and guidance rules"]
```

The evaluator uses the configured Cloud AI provider and must fail closed if that provider resolves to `peta`. It runs asynchronously through the existing AI job reliability controls and uses an evaluation-specific low-priority queue for backfills. Before the Cloud call, a deterministic transcript retriever selects bounded candidate spans for each verifiable unit; retrieval never searches another meeting or an external corpus.

The request contains only the transcript and the deterministic PropertyLab review units needed for the meeting. Its metadata includes run UUID, source generation UUID, source PropertyLab request ID, evaluator schema version, evaluator prompt SHA, and queue attempt.

## 6. Metric definitions

Dashboard denominators and exclusions are explicit:

- **Human review coverage:** saved section reviews ÷ eligible successful PropertyLab generation sections.
- **Average accuracy/completeness/actionability:** arithmetic mean of saved human scores; no imputation for missing reviews.
- **Quality pass rate:** reviewed sections where all three scores are at least 4 ÷ reviewed sections.
- **Human-confirmed unsupported rate:** saved reviews marked `Issue found` ÷ saved reviews marked either `Issue found` or `No issue`. `Not sure` and unanswered are excluded and displayed separately.
- **Automatic flag rate:** review units marked contradicted or no-evidence ÷ evaluated units typed transcript-verifiable. Judgment/advice and external-fact units are excluded from this denominator and reported separately. This is a diagnostic signal and is never labelled hallucination rate.
- **Claim confirmation rate:** claim evaluations with a human verdict ÷ claim evaluations shown for confirmation.
- **First-pass JSON parse rate:** PropertyLab attempts whose raw assistant response parsed as JSON before repair/normalization ÷ PropertyLab attempts with a provider response.
- **First-pass schema success rate:** PropertyLab attempts satisfying the strict canonical schema before defaults/coercion ÷ PropertyLab attempts with a provider response.
- **P95 full analysis time:** nearest-rank 95th percentile of generation `completed_at - queued_at`; provider duration is shown separately and P95 is marked unstable below 20 samples.
- **Failure rate:** settled failed PropertyLab generations ÷ all settled PropertyLab generations.
- **Memory stability:** trend of VRAM baseline and OOM/restart counters across comparable benchmark batches. “No leak” is reported only when pre/post idle VRAM returns within a configured tolerance and no monotonic growth is observed over at least three batches.

### 6.1 Cost

Show three values separately:

1. **Cloud evaluator vendor cost** from the evaluator `ai_requests` rows.
2. **PropertyLab request-equivalent GPU cost estimate** = configured GPU hourly rate × end-to-end request seconds ÷ 3600. This may double-count concurrent requests and must be labelled an estimate.
3. **Benchmark allocated GPU cost per meeting** = actual benchmark run wall-clock billable cost ÷ completed PropertyLab generations, with optional weighted allocation by request duration. This is the primary Phase 0 comparison figure because it preserves idle/loading/concurrency overhead inside the run window.

Disk, EIP, and network charges are reported separately when configured. The app must never claim an exact billing cost when a run has no verified billable window.

## 7. Training guidance rules

Guidance is deterministic, versioned, snapshotted, and advisory. Every card shows its sample size, threshold, evidence, and why the rule fired. Snapshots freeze filters, metrics, rule version, recommendation, and candidate case IDs so later prompt/model changes cannot rewrite historical decisions.

Ten evaluated meetings and 50% human section-review coverage are enough to show explicitly labelled pilot signals. A qualified recommendation requires at least 30 completed generations and 100 human-confirmed review units. Below that threshold, show `Collect more evidence` and never recommend fine-tuning.

### Prompt candidate

Recommend prompt/schema work when any is true:

- first-pass schema success is below 99%;
- repeated instruction/format categories exceed 5% of confirmed issues;
- average accuracy, completeness, or actionability is below 4/5 and the transcript contains the missing/ignored evidence; or
- the same provider/model improves materially under a newer prompt version.

### RAG candidate

Recommend RAG when at least 30 claim reviews have `verdict=problem` and human issue category `external_knowledge`, they comprise at least 60% of actionable confirmed errors, and a permission-filtered authoritative corpus has been declared available. RAG is not recommended merely because the model omitted transcript content.

### Fine-tuning candidate

Recommend fine-tuning only when all are true:

- at least 500 adjudicated development examples share the same stable behavioral error taxonomy and a separately held-out set contains at least 100 examples;
- first-pass schema success is at least 99%;
- at least two consecutive prompt versions have been tested and the issue persists at 10% or more;
- RAG is available for cases requiring external facts, or the issue does not require external facts;
- the training/holdout split can be created without placing the same meeting in both; and
- quality remains below the 4/5 target or the confirmed unsupported rate remains above 5%.

Blind/holdout rows are never eligible for training. Ten local meetings are sufficient for workflow validation, not for a qualified training decision.

Each generation has an explicit evaluation cohort (`development` or `holdout`). Assignment is meeting-grouped so generations from the same meeting can never appear in both cohorts. Where a numeric prompt version cannot be resolved, consecutive prompt versions are ordered by the first-seen time of distinct resolved prompt SHA-256 values.

## 8. Security and privacy

- All evaluation dashboard and run-control routes require authenticated super-admin access at both middleware and controller boundaries.
- Salespeople see only meetings already allowed by the existing Zoom/lead visibility rules.
- A claim decision is accepted only for the current PropertyLab generation and a visible meeting.
- Cloud evaluator calls reuse the existing provider secret configuration; no key is stored in a request, Vue prop, log, Git file, or database metadata.
- Metrics collection uses a fixed server-side private/TLS URL, a read-only credential, short timeouts, response size/content-type checks, and an allow-list parser.
- Aggregate pages exclude full transcripts and raw AI request bodies.
- Reviewer notes and evidence excerpts are case-detail data, not exported by default.
- Run creation is capped and idempotent. There is no automatic recurring backfill in the first release.
- Cloud evaluation is default-off and enforces an approved-provider allow-list, daily request/token/vendor-cost ceilings, maximum units, maximum evidence size, and maximum output tokens.
- Transcript and PropertyLab output are delimited and treated as untrusted prompt content; instructions contained inside them must be ignored.
- Daily evaluator budgets are checked both when a run is created and immediately before every Cloud call. Usage comes from `ai_requests` rows for the evaluator prompt key within the configured Kuala Lumpur billing day. Exhaustion settles work as explicitly skipped/budget-exhausted; it is never treated as success.

## 9. Non-goals

- automatic model training or deployment;
- automated Go/No-Go approval;
- modifying Alibaba Cloud instance state or billing;
- public vLLM or Prometheus endpoints;
- timestamp/speaker evidence before transcript ingestion preserves those fields;
- K8s, multi-GPU, multi-node, or Phase 1 architecture.

## 10. Acceptance criteria

- A salesperson can save `No issue`, `Issue found`, or `Not sure` without learning AI terminology.
- Each successful PropertyLab generation records immutable provider/model/prompt/schema/transcript lineage and raw-vs-normalized reliability metrics.
- Cloud AI can evaluate frozen review units asynchronously; invalid evidence cannot be stored as supported/contradicted.
- Human and automatic verdicts remain separate and auditable.
- A super-admin can filter/drill into quality, latency, reliability, evaluation, and cost data.
- The dashboard never reports automatic flags as human-confirmed unsupported content.
- Guidance explains Prompt/RAG/fine-tune criteria and returns insufficient-data when thresholds are unmet.
- Guidance snapshots remain reproducible after later model/prompt/rule changes.
- GPU metrics fail safely and display stale/unavailable honestly.
- Existing Cloud AI/PropertyLab comparison, per-user reviews, non-webinar scope, and stale-generation protection remain intact.

