# petaV3 Dual-Region Real Estate AI Engine: Cloud GPU Validation and Private Deployment Plan

Version: v2.1

Date: 2026-08-06

Audience: PropertyLab / petaV3 leadership, engineering, mentors, and potential pilot customers

Code baseline reviewed: `codex/xiaochengxu`, HEAD `05ee8ae22b18`

Scope: Current-code assessment, cloud GPU POC, CN / Global regional deployment, AI flywheel, and private customer delivery

> v2.1 adds an Alibaba Cloud International Phase 0 runbook, business outcome labels from the first POC week, a USD 500 cloud-spend hard gate, correct economical-stop controls, and model-image and TLS security requirements. Review input came from the material shared on 2026-08-06; cloud-product facts are governed by official Alibaba Cloud and vLLM documentation.

## 1. Executive Summary

PropertyLab does not need to train a foundation model from scratch. The practical product is a privately deployable **Real Estate AI Engine** that:

1. Reuses petaV3's existing AI gateway, business data, and real estate workflows.
2. Validates self-hosted open-weight models on cloud GPU before buying hardware.
3. Operates separate China and international data planes.
4. Uses Workers for high-frequency signal extraction and a regional Orchestrator for sales judgement and Next Best Action.
5. Records whether recommendations are adopted and connects them to appointments, viewings, bookings, wins, and losses.
6. Shares only approved aggregate metrics, scoring methods, and versioned method packages; raw customer data remains regional.

The requested “two servers structure” should be treated as two complete deployment regions, not two individual machines:

| Region | Users | Data retained in region | Regional AI |
|---|---|---|---|
| CN Data Plane | China companies, agents, and China-sourced customers | CRM, messages, calls, recordings, files, vectors, AI logs, business outcomes | Models and GPU services approved for the China deployment |
| Global Data Plane | International customers or customers contractually assigned to Global | WhatsApp / LINE / Zoom, CRM, files, vectors, AI logs, business outcomes | International model APIs or Global cloud GPU |

Both regions use the same petaV3 code, schema contract, prompt contract, and business method. Their databases, object storage, vector stores, keys, model endpoints, and operational logs remain separate.

### 1.1 Roadmap

| Stage | Main Goal | Deployment | Main Output |
|---|---|---|---|
| Test / Phase 0 | Prove petaV3 can use a self-hosted model and enforce regional boundaries | Global cloud GPU plus two isolated test environments | `internal_ai` provider, one Worker, one Orchestrator path, evaluation report |
| Phase 1 | Build the internal regional brain and measurable flywheel | PropertyLab cloud, region-isolated | RAG, recommendation events, adoption, outcome linkage, internal case study |
| Phase 2 | Deliver private AI to external real estate companies | Customer private cloud, on-prem GPU, or regional managed deployment | CN / Global delivery package, pilot, SLA, commercial offer, operations SOP |

### 1.2 Core Decisions

- Rent cloud GPU first; do not buy on-prem hardware yet.
- Integrate through the existing `AiClient`; do not create a second business AI SDK.
- Prioritize RAG, workflows, and evaluation over pretraining.
- Start with drafts and recommendations reviewed by humans.
- Keep raw data regional before designing any cross-region aggregation.
- Prefer isolated customer deployments for early Phase 2 customers; do not treat the current `group_id` as a legal tenant boundary.
- Use Alibaba Cloud International for Phase 0, with Kuala Lumpur `ap-southeast-3` first and Singapore `ap-southeast-1` only as an inventory or capability fallback.
- Alibaba Cloud Kuala Lumpur remains part of the Global Data Plane; it must not hold real customer data classified for a Mainland China data plane.
- Treat the auditable chain of input, AI recommendation, human action, and business outcome as the POC asset. The GPU itself is not the moat.

## 2. Code-Verified Baseline

### 2.1 Existing AI Foundation

| Capability | Current Implementation | Code Evidence | Assessment |
|---|---|---|---|
| Unified AI calls | `AiClient::prompt/chat/stream` | `src/Ai/Services/AiClient.php` | Suitable as the AI gateway |
| Provider adapters | Anthropic, OpenAI, Gemini, DeepSeek; Seedance and Deepgram as non-chat providers | `src/Ai/AiCredential.php`, `config/ai.php` | Provider catalog exists; `internal_ai` does not |
| OpenAI-compatible transport | OpenAI and DeepSeek share `OpenAiTransport` | `src/Ai/Transports/OpenAiTransport.php` | vLLM can reuse this transport |
| Prompt registry | Registered keys, admin overrides, versions, provider/model pins | `config/ai_prompts.php`, `src/Ai/AiPrompt.php` | More mature than assumed in v1.0 |
| Request audit | PROCESSING / SUCCESS / FAILED, input, output, tokens, cost, duration, lead, subject | `ai_requests`, `AiRequestRepository` | Strong technical observability base |
| Queue resilience | Dedicated queue, rate limit, circuit breaker, retry, no-overlap | `app/Jobs/Ai/AiJob.php` | Ready for an internal endpoint |
| Transcription audit | Gemini / Deepgram attempts also write to `ai_requests` | `LoggingTranscriber.php` | Shared visibility already exists |
| Conversation analysis | Call / F2F / Zoom share one normalized analyzer | `src/Conversation/ConversationAnalyzer.php` | Best first Worker entry point |
| AI management UI | Provider keys, prompts, and request logs are manageable | Manage Integrations controllers and Vue pages | Can be extended for the internal model |

### 2.2 Current AI Business Uses

| Business Capability | Current Path | Phase 0 Treatment |
|---|---|---|
| Call / F2F / Zoom analysis | `ConversationAnalyzer` → `AiClient` | First-priority Worker |
| WhatsApp AI draft / auto reply | `GenerateAiReply` → `AiClient` | Draft-only POC |
| Portal AI Conversations | Streaming `AiClient` | Chat and streaming validation |
| AI Debate | `DebateRunner` across providers | Keep external models initially |
| Amenity demand | Queued `AnalyzeAmenityDemand` | Second-wave pilot |
| Catalogue AI content | Queued `GenerateCatalogueContent` | Public property-content pilot |
| Lead intelligence | `LeadIntelService` | Requires privacy and regional controls |
| Floor-plan / owner outreach | `OwnerListingController` | Candidate generation test |
| Transcription | Separate service with shared logging | No forced localization in the first POC |
| Video / image generation | Separate provider helpers and pipelines | Outside first-round localization |

### 2.3 Existing Data Foundation

- `countries`, `projects.country_id`, `catalog_projects.country_id`, and `property_analyses.country_id` provide a market-country dimension.
- Each catalogue ingestion adapter declares both `providerCode()` and `countryIso2()`.
- `catalog_project_sources` owns provider-specific identity and raw payload while the canonical catalogue remains source-agnostic.
- Hong Kong reference adapters can read HK data through the petaV2 reference connection.
- `engagements` records each lead × project lifecycle, including won / lost, `lost_stage`, and timestamps.
- `bookings` records booking, SPA, leader review, cancellation reason, and value.
- WhatsApp messages contain direction and delivery/read/failure state. AI auto replies carry generated, trigger, provider, and model metadata.
- The mini-program branch stores controlled AI-conversation state and durable inquiries: the application constrains candidate projects, the model has no direct database tool, and inquiry capture keeps a deterministic fallback when AI is unavailable.

### 2.4 Critical Gaps

| Gap | Current State | Impact |
|---|---|---|
| `internal_ai` provider | Missing from provider constants and catalog | vLLM cannot be selected in the UI or prompt pins |
| RAG / vector store | No Qdrant, pgvector, or embedding pipeline | Models cannot retrieve private petaV3 knowledge |
| CN / Global regions | No `data_region`, `deployment_region`, or residency policy | Regional isolation cannot be demonstrated |
| True tenant boundary | Authorization is mainly all / group / team / own | Insufficient for shared-database external tenancy |
| Orchestrator | No common task delegation / verify / synthesize service | Existing AI features remain independent use cases |
| Advisor / method package | No aggregate review or versioned distribution mechanism | The global brain remains conceptual |
| Recommendation adoption | WhatsApp Use and Dismiss both clear the same `ai_draft` | Adoption cannot be measured |
| Recommendation-to-outcome link | `ai_requests` is not attributed to engagement / booking outcomes | Business uplift cannot be proven |
| Regional audit fields | `ai_requests` has no tenant or region | Regional audit evidence is incomplete |

## 3. Product Definition

### 3.1 Real Estate AI Engine Components

1. **Model Gateway**: one interface for external and self-hosted models.
2. **Regional Workers**: structured extraction from messages, calls, meetings, portal, and field signals.
3. **Regional Orchestrator**: combines facts per lead, verifies conflicts, updates the customer picture, and proposes Next Best Action.
4. **Private RAG**: retrieves customer, property, workflow, and knowledge documents inside the permission and regional boundary.
5. **Recommendation Sensor**: records shown, accepted, edited, sent, executed, dismissed, or expired actions.
6. **Outcome Join**: connects recommendations to appointments, engagements, bookings, SPA, wins, and losses.
7. **Method Control Plane**: versions prompts, scoring rules, workflows, and safety policies.
8. **Audit and Evaluation**: measures quality, latency, cost, errors, adoption, and uplift.

### 3.2 Non-Goals

- No foundation-model pretraining from scratch.
- No complete localization of video, image, and speech models during Test.
- No China raw customer data sent to international model APIs.
- No assumption that embeddings are anonymous and freely transferable.
- No claim of conversion uplift before adoption and outcome evidence exists.
- No shared-database multi-tenant offer before tenant redesign or isolated deployments.

## 4. Target Architecture

```text
                         PETA Method Control Plane
                  prompt / workflow / score / policy versions
                    aggregate benchmark / release registry
                         NO raw customer data
                          /              \
              signed method package    signed method package
                       /                    \
          CN Regional Data Plane       Global Regional Data Plane
        -------------------------      ---------------------------
        China channel adapters         WhatsApp / LINE / Zoom / CRM
        Regional API + queue            Regional API + queue
        Operational database            Operational database
        Private object storage          Private object storage
        Regional vector store           Regional vector store
        CN model gateway / GPU           Global model gateway / GPU
        Worker + Orchestrator            Worker + Orchestrator
        Regional AI audit logs          Regional AI audit logs
```

### 4.1 Meta LOOP Mapping

| Role | petaV3 Responsibility | Frequency | Data Location |
|---|---|---|---|
| Worker | Transcription, entities, intent, objections, budget, sentiment, drafts | High-frequency and parallel | Always regional |
| Orchestrator | Merge Worker results, conflict checks, lead state, NBA, escalation | Important event or schedule | Always regional |
| Advisor | Method review, scoring calibration, risk review, version suggestions | Low-frequency / monthly | Approved aggregate data only |

Model names are not architecture. Worker, Orchestrator, and Advisor require stable interfaces; the underlying model is selected by region, cost, quality, and compliance.

### 4.2 Meaning of “One Brain”

“One brain” must not mean one database containing all raw data. It should mean:

- One signal schema.
- One scoring and workflow contract.
- One versioned prompt / policy release mechanism.
- Two regions operating the same method independently.
- Controlled export of non-sensitive evaluation aggregates only.

### 4.3 Routing Rules

| Field | Meaning | Example |
|---|---|---|
| `market_country` | Property market | `HK` |
| `deployment_region` | Region where the application instance runs | `CN` |
| `data_region` | Region where a business record may be stored | `CN` |
| `tenant_region` | Customer company's primary data region | `CN` |
| `source_region` | Region/provider from which reference data originates | `GLOBAL` |

For a China agency selling Hong Kong property:

- The property catalogue has `market_country=HK`.
- The agency, WeCom customer, conversation, recording, and analysis remain `data_region=CN`.
- Licensed public property data may be published as a read-only projection into CN.
- China customer data must not move to Global because the property is in Hong Kong.

## 5. Data Contract Direction

This section defines proposal-level contracts. Exact columns, deletions, and migration order must remain aligned with the Database Redesign document.

### 5.1 Existing Tables to Reuse

| Table | Continued Responsibility |
|---|---|
| `countries` | Market country, currency, locale, timezone |
| `catalog_projects` | Cross-source canonical property projects |
| `catalog_project_sources` | Provider identity, fields, raw payload, scrape time |
| `projects` | Working projects used by business workflows |
| `leads` | Customer and lead entry |
| WhatsApp, call, F2F, Zoom tables | Regional interaction facts |
| `engagements` | Lead × project sales lifecycle |
| `bookings` | Booking / SPA / cancellation outcomes |
| `ai_requests` | Technical model-call audit |
| `ai_prompts`, `ai_prompt_versions` | Prompt content and versions |

### 5.2 Regional Contracts

For isolated customer deployments, `deployment_region` may be fixed by environment. A tenant table and row-level tenant key are required only before a future shared platform.

| Object | Proposed Field / Mechanism | Purpose |
|---|---|---|
| Deployment | `DEPLOYMENT_REGION=CN|GLOBAL` | Prevent calls to the wrong regional endpoint |
| AI request | `deployment_region`, optional `tenant_ref` | Audit and cost ownership |
| Data export | Export allowlist plus batch audit | Prove only approved aggregates cross regions |
| Catalogue projection | Source, licence, published region, version | Govern public product-data replication |
| Model endpoint | Region, provider, model, policy version | Prevent CN traffic from using Global models |

### 5.3 Flywheel Tables

#### `ai_recommendations`

Core fields: `uuid`, `ai_request_id`, `lead_id`, `engagement_id`, `conversation_id`, `recommendation_type`, `payload`, `deployment_region`, `prompt_version`, `model`, and `generated_at`.

#### `ai_recommendation_events`

Append-only events: generated, shown, accepted, edited, sent, executed, dismissed, expired; with recommendation, actor, message, metadata, and timestamp.

#### `ai_outcome_links`

Links a recommendation to engagement / booking outcomes such as appointment, booked, SPA signed, won, lost, or cancelled, including outcome time and attribution window.

Existing `engagements` and `bookings` remain the source of truth. The new tables only provide an auditable recommendation-to-outcome relationship.

### 5.4 Method Control Tables

Add in Phase 2:

- `method_packages`
- `method_package_versions`
- `method_package_deployments`
- `regional_export_batches`

A method package may contain prompts, JSON schemas, scoring parameters, workflow definitions, and compatible model lists. It must not contain raw customer examples.

### 5.5 Regional Database Principles

- No database-level foreign keys between CN and Global.
- No real-time cross-region joins.
- No replication of leads, messages, recordings, raw documents, or embeddings.
- Catalogue data moves only through an explicit publish/export workflow.
- Aggregates require de-identification and a minimum sample threshold.
- Method-package releases and aggregate exports record version, operator, hash, and time.

## 6. Test / Phase 0: Cloud GPU and Regional Feasibility

Recommended duration: four weeks.

Scope: one Global cloud GPU, two isolated test environments, synthetic or approved de-identified data.

Phase 0 validates deployment, integration, quality, cost, and regional controls. It does not perform foundation-model pretraining or distillation. Fine-tuning or distillation should be considered only after real evaluation proves a general model cannot meet the task target and a sufficient outcome-labeled dataset exists.

### 6.1 Week 0: Data and Baseline

1. Select three tasks: conversation analysis, WhatsApp draft, and lead Next Best Action.
2. Build representative, approved, de-identified samples.
3. Create human gold answers and a scoring rubric, and freeze a blind evaluation set before prompt tuning.
4. Label source, customer type, language, market country, permitted region, consent, and retention status.
5. From the first week, capture verifiable outcomes: appointment booked, viewing completed, booking, SPA signed, won/lost, event time, amount, label source, and evidence reference.
6. Record unconfirmed outcomes as `unknown`; never replace factual labels with model inference.
7. Establish current external-provider quality, latency, and cost baselines.

Acceptance:

- Target at least 100 conversation-analysis examples, with 20-50 dual-reviewed samples frozen as a blind evaluation set. If data is limited, 50 is the POC minimum and the limitation must be reported.
- Target at least 100 WhatsApp-draft examples and 50 lead/NBA scenarios.
- Known key fields and business outcome labels are at least 90% complete in the blind set; unknown outcomes are reported separately and excluded from uplift denominators.
- No unapproved identity document, bank account, full phone number, or raw recording.

### 6.2 Week 1: Cloud GPU

1. Use a company-owned Alibaba Cloud International account and a least-privilege RAM operator. If broad bootstrap access is temporarily required, narrow it immediately after provisioning.
2. Check live inventory and price in Kuala Lumpur `ap-southeast-3` first. Evaluate Singapore `ap-southeast-1` only when the required capability or inventory is unavailable, and record the reason.
3. Prefer `ecs.gn8is.2xlarge` (1 × NVIDIA L20 48GB, 8 vCPU, 64 GiB). If L20 is unavailable, use an A10 24GB-class fallback with Qwen3-14B AWQ. The creation page on the deployment date is authoritative for availability.
4. Use pay-as-you-go ECS, a VPC, 300-500GB ESSD, and disk encryption. Use an EIP only when cross-cloud connectivity requires it; prefer VPN or private connectivity.
5. Allow only the petaV3 source to reach port 443. Restrict SSH to a bastion, VPN, or fixed engineer IPs. Do not expose port 8000 to the internet or other VPC hosts.
6. Install the NVIDIA driver, Docker, and NVIDIA Container Toolkit.
7. Run a fixed vLLM tag and approved image digest. Record the model repository revision, tokenizer revision, and quantization checksum in a deployment manifest.
8. Run Qwen3-14B BF16 on L20 and a quality-validated 14B AWQ build on the A10 fallback. Structured extraction is the primary first-round decision task.
9. Bind vLLM to localhost and enable a dedicated service API key. Use a publicly trusted certificate or private mTLS at the gateway. Self-signed TLS combined with `verify=false` is prohibited.
10. Set a USD 500 POC cloud-spend hard cap with alerts at 50%, 80%, and 100%. Save evidence from the live calculator or order page.
11. Schedule non-working-hour stops with OOS or ECS API `StopInstance` using `StoppedMode=StopCharging`. Linux `shutdown` or `poweroff` does not enter economical mode and must not be used as a cost-control mechanism.
12. Record model version, quantization, GPU, VRAM, startup parameters, continuing EIP / disk charges, and actual daily cost.

Acceptance:

- `/v1/models` and `/v1/chat/completions` are reachable only from authorized petaV3 environments.
- JSON, streaming, timeout, and error behavior are repeatable.
- GPU, throughput, first-token latency, and memory are monitored.
- Budget alerts are tested and the instance enters economical mode outside working hours; the startup-capacity risk after economical stop is documented.
- TLS verification stays enabled, no client uses `verify=false`, and the deployment manifest contains no floating `latest` tag.

### 6.3 Week 2: petaV3 `internal_ai`

The current code can reuse `OpenAiTransport`, but the provider must be registered:

1. Add `PROVIDER_INTERNAL_AI` to `AiCredential`.
2. Add it to the provider catalog and chat-provider list.
3. Configure label, base URL, verify path, chat path, model catalog, and estimated cost.
4. Keep base URL in environment configuration and service token in encrypted `ai_credentials`.
5. Add the POC model ID to the admin-selectable catalog.
6. Verify `AiKeyService` model resolution and prompt pins.
7. Test `prompt()`, `chat()`, `stream()`, JSON mode, and failure logging.
8. Add provider, transport, model-validation, and request-log tests.

Acceptance:

- Manage AI Providers can store the internal endpoint token.
- Manage AI Prompts can pin a prompt to `internal_ai`.
- `ai_requests` records provider, model, prompt, tokens, duration, and errors.
- Existing external providers do not regress.

### 6.4 Week 3: Minimum Worker and Orchestrator

Worker:

1. Pin `conversation_analysis` to `internal_ai`.
2. Run the shared Call / F2F / Zoom schema.
3. Blind-compare with the current provider and calculate field-level accuracy for budget, objections, preferences, intent, and next action.

Orchestrator:

1. Register a prompt key such as `regional_orchestrator`.
2. Input only structured Worker output, lead facts, and the current engagement.
3. Return fixed JSON: temperature, buyer profile, lifecycle stage, risks, next action, confidence, and evidence.
4. Evidence must refer to regional facts.
5. Persist the result as a recommendation, not only response text.
6. Reserve `shown / accepted / edited / sent / dismissed / executed` events and link each recommendation to the Week 0 outcome labels.

Acceptance:

- One lead moves from transcript to Worker to Orchestrator.
- Output is schema-valid, auditable, and replayable.
- Failure never causes an automatic customer message.

### 6.5 Week 4: Regional Boundary Drill and Go / No-Go

Create isolated CN-test and Global-test environments with separate databases, storage, vector endpoints, model allowlists, and keys while using the same code and schema version.

Tests:

1. A CN synthetic tenant can write only to CN-test.
2. A Global synthetic tenant can write only to Global-test.
3. An approved HK catalogue projection can be published to CN-test.
4. CN lead, message, recording, and embedding data never appear in Global-test.
5. Only aggregate evaluation metrics reach the method control layer.

Minimum criteria:

| Metric | Suggested Threshold |
|---|---|
| Schema success | ≥ 95% |
| Human accuracy on critical fields | ≥ 85% and no worse than baseline |
| P95 non-streaming response | ≤ 15 seconds, task-specific |
| Failure rate | ≤ 2%, excluding deliberate fault tests |
| Regional-isolation tests | 100% pass |
| Severe data leakage | 0 |
| Cost per task | Explainable and compared with external providers |
| POC cloud spend | At or below the USD 500 hard cap; labor reported separately |
| Outcome labels | At least 90% complete for known outcomes in the blind set; unknowns reported separately |
| WhatsApp draft | Zero high-risk commitments; track edit rate, but edit rate above 50% is not an automatic No-Go |

Make Go / No-Go decisions by task. Whether Qwen3-14B is suitable for structured extraction is separate from whether it matches bilingual sales tone. Structured-extraction failure can block first integration; weaker draft quality may remain a human-review assist mode and should not invalidate the infrastructure POC by itself.

## 7. Phase 1: Internal Regional Brain and Flywheel

Recommended duration: six to ten weeks.

### 7.1 Data Scope

First wave:

- Lead profile and enrichment.
- WhatsApp conversations.
- Call / F2F / Zoom transcript and analysis.
- Engagement, appointment, and booking.
- Working project and published catalogue facts.
- Approved internal sales SOP.

Second wave:

- Portal behavior.
- Campaign and advertising signals.
- Badge / viewing field signals.
- Wealth planning and property analysis.

### 7.2 RAG Pipeline

1. Define owner, region, retention, and permission for every document type.
2. Read only from regional source databases or object storage.
3. Clean, chunk, and remove sensitive fields not needed for retrieval.
4. Generate embeddings inside the region.
5. Store them in the regional vector store.
6. Attach region, lead, group/tenant, document type, source ID, permission, and version metadata.
7. Apply permission filters before vector retrieval.
8. Preserve source citations.
9. Delete or rebuild vectors when source permissions or deletion state changes.
10. Test prompt-injection documents and isolate untrusted instructions.

### 7.3 Adoption Sensor

Use WhatsApp draft as the first complete path:

1. AI creates a recommendation.
2. UI display records `shown`.
3. Use records `accepted`; it must no longer share an untyped clear API with Dismiss.
4. Agent edits record `edited` and a minimal difference metric.
5. Actual delivery records `sent` and outbound message ID.
6. Dismiss records `dismissed`.
7. Unhandled suggestions record `expired`.

### 7.4 Outcome Feedback

1. Treat appointments, engagement status, booking, SPA, and won/lost as outcome events.
2. Use lead × project engagement as the main business relation.
3. Define attribution windows.
4. Retain non-adoption comparison groups.
5. Evaluate by language, market, team, task type, and model version.

### 7.5 Phase 1 Outputs

- Internal AI Brain release.
- Regional RAG with permission and deletion synchronization.
- Recommendation, adoption, and outcome chain.
- Technical quality dashboard.
- Adoption, conversion, and uplift dashboard.
- At least one internal case study.
- On-prem GPU Go / No-Go report.

## 8. Phase 2: External Customers and Production Regions

### 8.1 Delivery Models

| Model | Customer | Recommendation |
|---|---|---|
| Customer-isolated regional cloud | Early pilots and data-sensitive firms | Highest |
| PropertyLab-managed single tenant | Managed service with database isolation | High |
| Customer on-prem GPU | Customer has facilities, IT, and strict egress controls | Medium |
| Shared-database multi-tenant | Scaled SaaS | Defer until tenant redesign |

### 8.2 CN Production Prerequisites

1. Confirm China cloud, networking, domain, registration, and contracting requirements.
2. Obtain professional review of PIPL, data export, and recording processing.
3. Select models or GPU services available in-region and permitted by customer policy.
4. Implement China channel adapters such as WeCom.
5. Establish CN object storage, vector store, backup, KMS, and monitoring.
6. Deny or allowlist egress to Global endpoints.
7. Implement approval and audit for aggregate exports.

### 8.3 China Agency Selling Hong Kong Property Pilot

1. Maintain HK canonical projects in the Global catalogue.
2. Publish an approved sales projection into CN.
3. China agents engage through the CN CRM and channels.
4. CN Workers analyze messages, calls, and meetings.
5. The CN Orchestrator proposes matching, objection handling, and NBA.
6. Agents accept, edit, or reject recommendations.
7. Appointments, bookings, and lost reasons remain in CN.
8. Only thresholded anonymous aggregates enter method review.

### 8.4 Pilot Delivery Steps

1. Data-processing and security discovery.
2. Data inventory and regional classification.
3. Select two or three high-value use cases.
4. Deploy the isolated environment.
5. Configure channel adapters and SSO / roles.
6. Import approved knowledge and catalogue data.
7. Run shadow mode.
8. Run a human-in-the-loop pilot.
9. Evaluate adoption and outcomes.
10. Complete production review, SLA, and support handover.

## 9. Model and GPU Strategy

| Layer | Task | Strategy |
|---|---|---|
| Worker | Extraction, classification, short summary, draft | Small, fast, regional, low cost |
| Orchestrator | Synthesis, conflict checking, NBA | Medium, high-quality model |
| Advisor | Monthly review and strategy critique | Low-frequency premium model or human + model |

The Phase 0 reference combination is Alibaba Cloud International Kuala Lumpur `ecs.gn8is.2xlarge`, one L20 48GB GPU, and Qwen3-14B BF16. If live inventory does not support it, evaluate the same class in Singapore, then an A10 24GB plus Qwen3-14B AWQ. Alibaba Cloud documents that gn8is is available only in selected regions and zones, so this plan does not promise stock in a fixed region.

The POC is not expected to choose a permanent model. Qwen3-14B first proves structured extraction, JSON schema stability, and cost. Sales tone and complex bilingual drafts remain benchmarked against the current frontier provider. An internal model that needs substantial editing may still be a Conditional Go when it avoids risky commitments and saves analysis time.

Record model revision, quantization, context, concurrency, tokens per second, P50/P95 latency, GPU utilization, memory, and cost per task.

Buy on-prem hardware only when Phase 0 and 1 pass, workload is stable, 12-24 month TCO is favorable, facilities and operations exist, and local deployment is a contract or compliance requirement.

## 10. Evaluation

### 10.1 Offline Quality

- JSON/schema validity.
- Human factual accuracy.
- Hallucination rate.
- Evidence coverage.
- Multilingual quality.
- Safety violation rate.
- Blind comparison with the current provider.
- Field-level precision / recall / F1 for structured extraction.
- Edit distance / edit rate and risky-commitment count for WhatsApp drafts; a single subjective overall score is not the only gate.

### 10.2 Online Technical Metrics

- Success / failure.
- Queue wait.
- First-token / total latency.
- Input / output tokens.
- Cost per task.
- GPU utilization.
- Fallback rate.
- Regional-routing violations.

### 10.3 Business Flywheel Metrics

- Recommendation shown rate.
- Adoption and edit rate.
- Sent / executed rate.
- Appointment rate.
- Booking / SPA / won rate.
- Loss and cancellation reason.
- Uplift between adopted and non-adopted groups.

Static answer quality is not sufficient commercial evidence. The north-star question is whether recommendations improve valid business outcomes under an explainable comparison.

## 11. Security and Compliance

- Use private networking, TLS, service authentication, and allowlists.
- Continue using encrypted petaV3 credentials for API keys.
- Treat messages, recordings, files, embeddings, prompt logs, and model outputs as sensitive.
- Define retention, redaction, and least privilege for `ai_requests` request and raw response data.
- RAG applies authorization before retrieval.
- Each region owns backup, KMS, incident logs, and restore tests.
- Prohibit testing real customer data in personal AI accounts.
- Sign, version, and audit method packages and aggregate exports.
- Professional legal review is required; this plan is not legal advice.

## 12. Operations

Each region requires petaV3 application/API, web and queue workers, operational database, Redis, private object storage, vector store, model gateway, monitoring, centralized logs, and alerting.

Failure rules:

- CN must never automatically fall back to a Global provider.
- Global fallback is controlled per prompt policy.
- Interactive chat fails soft.
- Queued jobs use existing retry and circuit-breaker behavior.
- Recommendation failure never blocks core CRM work.
- Automatic customer messaging stays disabled until specifically approved.

Release process:

1. Build regional images from the same commit.
2. Apply region-specific configuration and secrets.
3. Deploy shadow / canary.
4. Run regional-boundary tests.
5. Run prompt regression.
6. Check schema and model compatibility.
7. Release production traffic.

## 13. Commercial Model

Sellable components:

- AI readiness assessment.
- Private Real Estate AI Engine deployment.
- Channel / CRM integration.
- RAG knowledge setup.
- Sales conversation intelligence.
- Next Best Action and agent copilot.
- AI governance, audit, and evaluation dashboard.
- Managed model / GPU operations.
- Quarterly method-package updates.

| Fee | Description |
|---|---|
| Discovery | Data, compliance, use-case, and infrastructure assessment |
| Implementation | Deployment, integration, migration, configuration, acceptance |
| Infrastructure | Customer direct cost or pass-through |
| Monthly platform/support | Operations, monitoring, upgrade, backup, support |
| Usage | Tokens, GPU, ASR, storage, or task volume |
| Method package | Versioned real estate methods and evaluation updates |

Avoid pure performance pricing before reliable uplift evidence exists. A capped, precisely defined performance component may be added after a successful pilot.

## 14. Team and Responsibilities

| Role | Responsibility |
|---|---|
| Product / Business Owner | Use cases, success metrics, pilot customer, commercial package |
| petaV3 Engineer | Provider integration, events, RAG, regional routing, tests |
| AI / ML Engineer | Model serving, evaluation, prompt/model experiments |
| DevOps / Security | Cloud, network, secrets, backup, monitoring, incident response |
| Data Owner | Authorization, quality, retention, gold dataset |
| Legal / Compliance | China and international data-processing review |
| Sales / Customer Success | Discovery, training, adoption, case study |

## 15. Timeline

| Time | Milestone |
|---|---|
| Week 0 | Data inventory, gold dataset, baseline |
| Week 1 | Global GPU on Alibaba Cloud International, Kuala Lumpur first, plus vLLM |
| Week 2 | `internal_ai` provider |
| Week 3 | Worker and Orchestrator POC |
| Week 4 | Regional drill and Go / No-Go |
| Month 2-3 | RAG, adoption, outcomes, internal case study |
| Month 4-6 | CN / Global external pilot |
| After pilot | Decide on local GPU and product scale |

Phase 0 has a USD 500 cloud-spend hard cap, not a long-term monthly quote. Recalculate from the Alibaba Cloud calculator or order page on the provisioning date, including model choice, runtime, disk, snapshots, EIP, and traffic. Engineering and labeling labor must be reported separately.

## 16. Risks and Mitigation

| Risk | Mitigation |
|---|---|
| Treating market country as data region | Separate market country, tenant region, and data region |
| Treating `group_id` as tenant | Isolated early deployments; tenant redesign before shared SaaS |
| CN data calls a Global model | Region policy, egress deny, integration tests |
| RAG leaks another customer's data | Physical isolation or strict tenant filtering and permission-first retrieval |
| Copilot without a flywheel | Require recommendation, adoption, and outcome in Phase 1 |
| False attribution | Attribution windows, comparison groups, segmented analysis |
| Model quality drift | Gold dataset, prompt/model versions, regression suite |
| GPU cost overrun | Queues, batching, concurrency caps, auto-shutdown, task-cost dashboard |
| Linux shutdown but charges continue | Use only OOS or ECS API `StoppedMode=StopCharging`, then verify state and billing |
| Instance cannot restart after economical stop | Keep schedule buffer, pre-start inventory checks, and a Singapore / alternate-spec runbook |
| Alibaba Cloud Malaysia is mistaken for a China region | State that `ap-southeast-3` is Global; build a separate Mainland deployment with compliance approval for real China data |
| Automatic replies harm relationships | Draft first, human review, scenario-by-scenario release |
| Premature hardware purchase | Decide from measured concurrency and TCO |

## 17. Immediate Actions

### This Week

1. Approve the v2.1 architecture, regional definitions, and the rule that Kuala Lumpur is Global only.
2. Complete the AI Data Readiness Audit.
3. Select three POC tasks, gold-dataset owners, and a business-outcome label owner.
4. Create the `internal_ai` provider engineering ticket.
5. Define recommendation, event, and outcome schemas.
6. Create the Alibaba Cloud company account, least-privilege RAM role, USD 500 budget, and alerts.
7. Capture live Kuala Lumpur L20 inventory and fallback pricing evidence.
8. Confirm which China channels are included in the first pilot.

### Within Two Weeks

1. Start cloud GPU and vLLM.
2. Complete petaV3 provider integration and tests.
3. Run the conversation-analysis baseline.
4. Build CN-test / Global-test boundary tests.
5. Produce the first quality, latency, and cost report.

### Within One Month

1. Complete the Worker-to-Orchestrator POC.
2. Complete regional data-leakage tests.
3. Hold the Go / No-Go review.
4. Define the Phase 1 RAG and flywheel backlog.
5. Prepare the China-agency / Hong Kong-property demo script.

## 18. Conclusion

petaV3 already has a more mature AI gateway, prompt management system, audit log, queue resilience, multi-country catalogue, and sales-outcome data than v1.0 assumed. Phase 0 should therefore register `internal_ai`, deploy cloud GPU, establish regional boundaries, and prove one Worker-to-Orchestrator path instead of rebuilding the AI calling layer.

The moat is not “we have a local model.” It is:

- A real estate-specific data contract.
- Regional multi-touch Workers.
- Auditable Next Best Action.
- Recommendation adoption joined to business outcomes.
- CN / Global raw-data isolation.
- Versioned, distributable, and measurable methods.

Cloud GPU is the validation instrument. The Real Estate AI Engine, data flywheel, and repeatable private deployment capability are the product.

## 19. Official References

- Alibaba Cloud region IDs: Malaysia (Kuala Lumpur) `ap-southeast-3` and Singapore `ap-southeast-1`: <https://www.alibabacloud.com/help/en/user-center/developer-reference/common-region-id-reference>
- Alibaba Cloud GPU instance families, including gn8is / L20 48GB specifications and regional-availability notes: <https://www.alibabacloud.com/help/en/ecs/user-guide/gpu-accelerated-compute-optimized-and-vgpu-accelerated-instance-families-1>
- Alibaba Cloud ECS economical mode, billing boundaries, and `StopInstance` requirements: <https://www.alibabacloud.com/help/en/ecs/user-guide/economical-mode>
- vLLM OpenAI-compatible server and API-key support: <https://docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server/>
- vLLM official Docker deployment guidance: <https://docs.vllm.ai/en/latest/deployment/docker/>
