# AI extraction validation and evaluation

## What it does

The extraction boundary accepts unconfirmed customer statements supported by an authorized source. The provider cannot set readiness, stage, priority, ownership, actions, blocks or contact policy. A manager must review facts before they affect readiness. A model opt-out signal is advisory; the established deterministic opt-out policy or an explicit human decision owns suppression.

This boundary is deliberately conservative. It checks authority, source identity, quotations, typed values, references and obvious unsupported interpretations. It cannot prove a model understood every multilingual sentence. The real-recording evaluation remains a separate release gate.

## How it works

`ExtractionValidator::validate($output, $packet, $confirmed)` returns:

- `proposals`: valid, confidence-ranked candidates (maximum 12). Each retains dimension, bare field key, typed value, quote and confidence, and adds `qualified_key`, a server-derived `source_span` and `conflicts_with_id`.
- `discards`: input index, qualified field when known, and a stable reason. No customer quote or raw sensitive value is needed for aggregate discard logs.
- `summary`, `intent`, `opt_out_detected`: advisory content; none is effective evidence or a command.
- `validator_version`: the validation contract used for this run.

### Trusted source packet

The caller supplies current, authorized `lead_id`, `scope_key`, `source_event_id`, `source_revision`, `observed_at` and `segments`. Each segment has a stable string `id`, `text`, `speaker_role`, `lead_id`, `scope_key`, `source_event_id`, and optional recording `start_sec`/`end_sec`. Context messages may use `context_only: true` but cannot support a proposal in the current source run.

Only an explicitly customer-attributed segment matching the packet's customer, intent and source is eligible. Unknown recording speakers remain unknown; transcript labels such as Speaker A/B do not prove customer identity. Quotations must match exactly after whitespace/case normalization. Punctuation, numbers and negation are retained. Ambiguous duplicate quotation matches require a segment ID. `verifyQuote()` rechecks a stored source span against a freshly authorized packet at read/review time; the calling repository still checks live source revision, access and current context.

### Typed values and references

The existing `FieldRegistry` remains authoritative for dimension, field, type and AI eligibility. Unknown keys, model actor/tenant/state fields, specialist assessments and measured system fields are rejected. Confidence must be a finite number at least 0.70.

Reference IDs require `packet.reference_ids`, keyed by singular reference name (`pool_id`, `party_id`, `source_event_id`, and so on). Current versions require `packet.versions`. Cash amount/reserve/availability require an explicitly selected `packet.bindings[qualified_key].pool_id`; knowing that several pools exist does not authorize choosing one. Route/target-dependent agreements require explicit current source/target/version bindings. The model never creates a funding pool, party, route, target or availability tranche. Missing references leave the information in the advisory summary for staff to clarify.

Party proposals require a known party label/alias in the quote and an explicit role phrase. Text fields preserve quoted text. Money must be present as a numeric literal, common MYR shorthand (for example RM80k or RM80 ribu), or a supported Chinese numeral (for example 八万). Dates must be explicitly quoted, and timing anchors retain the original observation day. Unsupported or ambiguous dates are discarded. Obvious negations, conditionals, reported third-party assertions and embedded instruction attempts are rejected for structured positive claims. These checks prioritize safe rejection over speculative recall; they are not a complete natural-language inference engine.

Accepted effective facts supplied for deduplication include `key`, `id`, `value`, `lead_id`, `scope_key`. Equivalent monetary spelling and set order do not produce new proposals. A different value retains the exact conflicting assertion ID. The repository must recheck that ID at save/review time and never replace confirmed evidence silently.

### Identity-number redaction

`ExtractionRedactor::text()` masks compact, hyphenated, spaced, digit-by-digit and Unicode-digit twelve-digit identity numbers. `packet()` applies it to free-text leaves and nested text lists while preserving metadata hashes and IDs. Invocation must redact before the provider call and before provider request logging. Do not save raw provider input in a second log. Identity status/consent remain secure system references, never inferred from spoken identity digits.

### First purchase inherits customer evidence without copying it

Creating the customer's first purchase attaches the existing intake when available. Current eligible Need and Relationship evidence is referenced in its original intake: confirmed evidence uses `EvidenceLink` with the exact review ID; unreviewed evidence uses one immutable `first_purchase_intake.pending_linked` journal event, because the link table requires a real review. No assertion, confirmation, ownership or next-action agreement is manufactured.

Every read rechecks access to the original intake, current source revision/access, original AI generation scope, and the pinned review decision. A changed review or revoked source removes the inherited assertion from effective readiness. Pending evidence remains pending and must be reviewed in the original intake; that later review does not silently upgrade an earlier pending association. The workspace links back to the original intake and disables cross-context editing. Later purchases, including purchases after a lost or archived first purchase, do not inherit this evidence. Further intentional relevance changes require a separate explicit workflow; this step does not add a relinking command.

### Evaluation gate

Run the offline evaluator with a private JSON file:

```
php scripts/revenue-journey/evaluate-extraction.php /private/labelled-results.json
```

It boots no application/database, calls no provider and changes no release flags. Input is one case per recording:

```
{
  "recording_id": "private-recording-reference",
  "provenance": "real_recording",
  "human_labelled": true,
  "labelled_by": "reviewer-reference",
  "labelled_at": "ISO timestamp",
  "languages": ["en", "zh"],
  "gold": [{"key": "need.purpose", "value": {"v": "investment"}}],
  "predicted": [{"key": "need.purpose", "value": {"v": "investment"}}],
  "raw_proposal_count": 1,
  "quote_failure_count": 0
}
```

The evaluator reports per-field and family precision/recall. The gate requires at least 20 distinct human-labelled real recordings, English/Chinese/Malay representation, positive labelled examples in Need/Cash/Decision-maker, precision at least 0.85 in those families and every observed field, and quote-verification failures strictly below 5% of raw model proposals. Repeated excerpts cannot inflate recording count. Labels must be created independently of model predictions; staff Accept/Reject clicks are review telemetry, not gold labels. The report trusts supplied provenance metadata, so a manager must verify the underlying labelled corpus before using it for release approval. A sparse corpus passing arithmetic is not evidence for unrepresented fields; report field coverage with the result.

Synthetic unit fixtures prove validator and metric behavior only. They do not satisfy this real-data gate. Keep wider proposal visibility and auto-drafts disabled until the independently reviewed evaluation is complete. ASR provider comparisons need a separate word-error-rate study on the same mixed-language corpus; this evaluator does not measure transcription quality.

## Related files

- `src/RevenueJourney/Ai/ExtractionValidator.php`
- `src/RevenueJourney/Ai/ExtractionValueGuard.php`
- `src/RevenueJourney/Ai/ExtractionRedactor.php`
- `src/RevenueJourney/Ai/ExtractionSchema.php`
- `src/RevenueJourney/Ai/ExtractionEvaluation.php`
- `scripts/revenue-journey/evaluate-extraction.php`
- `tests/Unit/RevenueJourney/Ai/ExtractionValidatorTest.php`
- `tests/Unit/RevenueJourney/Ai/ExtractionEvaluationTest.php`
- `src/RevenueJourney/Support/FieldRegistry.php`
- `config/readiness.php`
- `src/RevenueJourney/Repositories/FirstPurchaseEvidenceRepository.php`
- `tests/Feature/RevenueJourney/AiPreparationTest.php`
