# PetaHK Project Catalogue / Working Project Handoff

> Audience: the next Claude/Codex engineer taking over PetaHK project-data work.
>
> Date prepared: 2026-07-12 (Asia/Kuala_Lumpur)
>
> Implementation baseline inspected: `origin/master` at `80dac964` (`Update WhatSapp Broadcast Bug`, 2026-07-12 00:08 +08:00).

## 1. Why this document exists

Lee Jie split the PetaHK/New Project work into SSO, project list, project detail, project data, and eventual reuse of the new public project-detail structure inside the logged-in portal. You Chen's main ownership is **Project data**, with shared responsibility for **Project detail page**.

The database discussion established that PetaV3 needs two different concepts:

1. A global catalogue containing every scraped development across countries and across new/subsale markets.
2. A working project selected from that catalogue for operational workflows such as owner-listing imports, ad creation, brochures, campaigns, and outreach.

This document records the Mentor's decisions, translates them into technical invariants, and explains how the latest `master` currently differs from the intended design.

## 2. Product context

The long-term project-detail information architecture should follow the public Investhink project page pattern, for both public visitors and logged-in users:

- hero/gallery and project identity;
- location, developer, tenure, completion/key date, unit count and price range;
- project overview and facilities;
- map and nearby amenities;
- unit/floor-plan selector with size, price and plan image;
- rental, market value, PSF, ROI and financial projections;
- nearby supply/developer history;
- consultant/contact actions;
- logged-in-only analysis/actions layered on top of the same catalogue record.

The existing PetaHK prototype screenshots show 513 Hong Kong developments, a project/map list, a layout-analysis table, and a detail page with floor-plan-level rent/ROI/market-value information. These screens are catalogue consumers. They are not working-project management screens.

## 3. Mentor statements and confirmed decisions

### 3.1 Initial two-level model

Mentor described the intended concepts as:

- `project_catalogues`: all scraped projects, across all countries and both new/subsale markets, linked to related tables such as floor plans;
- `project`: a shortlisted/working project used for actions such as importing owner listings or creating ads.

The names were initially described as reference names. The latest code currently uses the physical table/model names `catalog_projects` / `CatalogProject` and `projects` / `Project`. This handoff treats `catalog_projects` as the current implementation name for the Mentor's logical `project_catalogues` concept, avoiding another cosmetic rename unless the team explicitly requests it.

### 3.2 Cross-source duplicates

Question: if the same real-world project is scraped from multiple websites, should each source create a separate catalogue project?

Mentor's answer:

- records should be merged;
- matching must be handled case by case;
- different names from different sources may require an explicit mapping;
- when sources disagree, the preferred source may differ by data type/field.

Technical meaning:

- one canonical catalogue row represents one agreed real-world development/phase;
- multiple source records may point to that canonical row;
- source identity and raw/source values must remain traceable;
- the canonical table cannot itself use `(data_provider_id, external_id)` as its identity because a canonical row may have many providers/external IDs;
- unsafe fuzzy matching must not silently merge records. Start with provider/external-ID identity, explicit mapping, and deterministic normalized-name candidates for admin review.

### 3.3 Tenant/company separation

Question: should working projects be separated by tenant/company?

Mentor's answer:

- no tenant layer is required;
- the current `app.propertylab.com.my` deployment is for Wai Kit's usage;
- another company/client such as Hong Kong receives a cloned deployment/repository/database such as `hk.propertylab.com.my`;
- catalogue data can be copied to that deployment when required;
- the solution should support future catalogue synchronization rather than assuming a permanent one-time copy.

Technical meaning:

- do not introduce `tenant_id`, tenancy middleware, tenant-aware repositories, or shared-database tenant partitioning;
- the catalogue is global **within one deployment**;
- current `group_id` behavior is internal authorization/workflow partitioning, not the tenant architecture rejected by Mentor; do not remove it as part of this work;
- cross-deployment synchronization should use stable catalogue UUIDs and idempotent import/export semantics rather than copying numeric primary keys blindly.

### 3.4 Working-project overrides

Question: may a working project change catalogue-derived values such as name, floor plans, pricing, rent or market value?

Mentor's answer:

- a working project cannot override catalogue data because it references `project_catalogues`;
- an admin may change catalogue data;
- only selected data should be manually overridable; examples given were rental price and market value.

Technical meaning:

- `projects` contains workflow identity/state and a required catalogue reference;
- catalogue-owned fields must not be editable from a working-project form;
- catalogue edits are global within the deployment and affect every consumer and every linked working project;
- manual overrides must be field-whitelisted;
- an active manual override must survive future scraper synchronization;
- removing an override should recompute the field from source data using the normal merge priority.

## 4. Terminology for the implementation

### Canonical catalogue project

Current code name: `catalog_projects` / `Src\Analysis\Reference\CatalogProject`.

One row per agreed real-world development/phase in one deployment. It owns normalized project facts and is the record read by PetaHK list/detail/analysis surfaces.

### Catalogue source record

Proposed code name: `catalog_project_sources` / `CatalogProjectSource`.

One provider-specific scraped record. It owns `data_provider_id`, `external_id`, source name/URL, raw payload, normalized source fields, and scrape timestamp. Many source records may belong to one canonical catalogue project.

Linking two differently named source records to the same canonical project is the explicit name/source mapping Mentor described.

### Catalogue floor plan

Proposed code name: `catalog_floor_plans` / `CatalogFloorPlan`.

A floor plan/unit type owned by a canonical catalogue project. Rent and market value are normally floor-plan-level values in the PetaHK UI, so the initial override design must support these fields here rather than treating them only as project-wide values.

### Working project

Current code name: `projects` / `Src\Property\Project`.

A project selected for an operational workflow. It keeps workflow fields such as UUID, group ownership, status, blame/timestamps, and the required `catalog_project_id`. It must not be an independent editable copy of catalogue facts.

### Workflow selection

Current code name: `focus_projects` / `Src\Property\FocusProject`.

Links a working project to New Project Suite or Subsale Suite. This remains useful unless a separate product decision removes it.

## 5. Target data flow

```text
Provider scrapers / source databases
        |
        v
catalog_project_sources  -- many source rows, preserved provenance
        |
        v
deterministic merge resolver
  1. active admin override
  2. configured provider priority for this field
  3. newest non-null source value
        |
        v
catalog_projects  -- one canonical development/phase
        |
        +--> catalog_floor_plans
        +--> PetaHK public list/detail/layout analysis
        +--> logged-in project detail and analysis
        |
        v
projects  -- operational selection, no catalogue-field overrides
        |
        +--> focus_projects (New Project Suite / Subsale Suite)
        +--> owner listing imports/rows/outreach
        +--> brochures, creatives, ads and campaigns
```

## 6. Current `origin/master` state

The latest `master` already contains a useful first foundation. Do not restart from the June schema and do not create parallel concepts without reading these files.

### 6.1 What already exists

- `database/migrations/2026_07_11_100001_rename_flg_projects_to_projects.php`
  - renames `flg_projects` to `projects`;
  - renames workflow types to `new_project_suite` and `subsale_suite`.
- `database/migrations/2026_07_12_100001_create_countries_table.php`
  - creates country/currency/locale/timezone support.
- `database/migrations/2026_07_12_100002_create_data_providers_table.php`
  - creates provider lookup data and seeds EdgeProp.
- `database/migrations/2026_07_12_100003_rename_edgeprop_projects_to_catalog_projects.php`
  - renames the local EdgeProp mirror to `catalog_projects`;
  - adds provider and country columns;
  - renames EdgeProp IDs to generic external IDs.
- `database/migrations/2026_07_12_100004_add_origin_and_country_to_projects_table.php`
  - renames the project link to `catalog_project_id`;
  - adds project origin and country.
- `database/migrations/2026_07_12_100007_add_hong_kong_to_countries.php`
  - adds Hong Kong country metadata.
- `src/Analysis/Reference/CatalogProject.php`
  - represents the current catalogue row and contains analysis helpers.
- `src/Property/Project.php`
  - links to `CatalogProject` and continues to own operational relations.
- `app/Http/Controllers/Manage/Property/CatalogController.php`
  - provides catalogue search and derived-layout JSON endpoints.
- `app/Console/Commands/ImportEdgepropData.php`
  - imports EdgeProp JSON into `catalog_projects`.
- `app/Console/Commands/ImportMarketData.php`
  - imports reference data from PetaV2 into local tables.
- `src/Property/Repositories/ProjectRepository.php`
  - seeds working-project floor plans and market anchors from a selected catalogue project.
- `tests/Feature/FacebookLeadGenerator/ProjectOriginTest.php`
  - locks current custom-vs-catalog behavior.

### 6.2 Where the current implementation conflicts with Mentor's final decisions

| Area | Current `master` | Required direction |
|---|---|---|
| Canonical identity | `catalog_projects` has one `data_provider_id` and one `external_id`; unique key is provider + external ID | One canonical row may have many provider source rows |
| Duplicate merging | Importers upsert directly by provider/external ID | Ingest a source record, map/merge it into one canonical row |
| Source provenance | Canonical row stores source identity and resolved values together | Preserve provider-specific source rows separately |
| Working fields | `projects` stores editable name, developer, location, prices and highlights | Working project must read catalogue facts and cannot override them |
| Custom projects | Working projects may be created with `origin=custom` and no catalogue link | Admin creates/edits a catalogue row first; working projects always reference one |
| Floor plans | `seedFloorPlansFromCatalog()` copies derived floor plans into `flg_floor_plans` per working project | Catalogue owns floor plans; working project references them |
| Admin catalogue editing | Catalogue controller exposes search/layout only | Admin needs whitelist-controlled catalogue editing |
| Manual override safety | `market:import` and `edgeprop:import` overwrite canonical fields | Active manual override must be sticky across sync |
| Deployment sync | No catalogue export/import contract | Stable UUID-based, idempotent package sync is required |

### 6.3 Important nuance: `group_id` is not tenancy

Current `projects.group_id` and `GroupScope` protect internal working data and allow sharing rules inside one deployment. Mentor rejected a multi-company tenant platform, not the existing group permission model. The catalogue should not have `group_id`, while working workflows may continue to use it.

## 7. Merge policy

Use this deterministic priority for every canonical field:

1. If the field is in the catalogue row's active `manual_overrides`, keep the admin value.
2. Otherwise, inspect all linked source records using the field-specific provider priority in `config/project_catalogue.php`.
3. Use the first non-null value from the highest-priority available provider.
4. If no provider has an explicit priority for that field, use the newest non-null source value.
5. Never overwrite a non-null canonical field with null.

Initial safe matching policy:

- exact `(data_provider_id, external_id)` always identifies the same source record;
- an existing source-to-canonical link always wins;
- normalized `country + project name + phase` may create a merge candidate, not an automatic destructive merge;
- an admin can link/reassign a source record to a canonical row;
- no fuzzy-distance auto-merge in the first implementation.

This satisfies the case-by-case requirement without silently combining similarly named but different Hong Kong phases/towers.

## 8. Manual override policy

Manual edits live on catalogue records, never on working projects.

The first whitelist should cover the examples explicitly mentioned by Mentor:

- catalogue/project-level market anchor fields already present in the current schema: `price_median`, `rental_yield`;
- floor-plan-level business values required by the PetaHK screens: `rental_price`, `market_value`.

The whitelist belongs in configuration so adding another approved field is a deliberate reviewable code change. Do not accept arbitrary request keys and do not expose raw JSON editing.

Each manual override needs metadata:

- field name;
- admin ID;
- override timestamp.

The resolved value remains in the normal typed column used by application queries. `manual_overrides` only records protection/provenance. Resetting an override removes its protection metadata and immediately reruns the merge resolver for that field.

## 9. Working-project policy

After migration:

- every `projects` row has a valid `catalog_project_id`;
- creating a working project means selecting a catalogue project and a workflow;
- project name/developer/location/pricing/floor plans are displayed from the catalogue relationship;
- working-project edit actions only change workflow-owned state;
- existing custom working projects are backfilled into new canonical catalogue rows before the catalogue link becomes required;
- brochures, campaigns, owner listing imports, owner rows and outreach remain attached to the working project;
- `focus_projects` remains the workflow selector.

For a low-risk migration, duplicated `projects` columns may temporarily remain as system-maintained compatibility caches, but no UI or request may edit them. The final removal should happen only after every consumer reads `catalogProject`.

## 10. Cross-deployment synchronization policy

PetaV3 and PetaHK are separate deployments/databases. The first sync mechanism should therefore be a versioned JSON export/import command rather than a tenant subsystem or a network service.

Required properties:

- stable catalogue UUIDs, not environment-specific numeric IDs;
- idempotent upsert by UUID;
- include canonical projects, source mappings, catalogue floor plans and active manual override metadata;
- preserve destination working projects by resolving catalogue UUIDs to local IDs;
- do not export `projects`, `focus_projects`, campaigns, leads, owner rows or other deployment-specific workflow data;
- a dry-run summary before writes;
- transaction per catalogue aggregate;
- reject unsupported package versions.

Media binary replication is outside this data-foundation plan. Only stable media/document URLs already stored as catalogue values may travel in the package. A separate asset-copy design is required before local-file media can be synchronized safely.

## 11. Scope boundaries

This handoff and the linked implementation plan cover:

- canonical multi-source catalogue data;
- deterministic merging and source mapping;
- admin whitelist overrides with scraper protection;
- reference-only working projects;
- catalogue-owned floor plans;
- command-based future synchronization between deployments.

They do not cover:

- SSO;
- a multi-tenant database architecture;
- the full PetaHK public project-list UI;
- the complete Investhink-style project-detail UI;
- fuzzy/AI automatic duplicate merging;
- copying local media binaries between deployments.

The list/detail pages should consume the resulting catalogue APIs after this foundation is stable.

## 12. Acceptance criteria

The data foundation is complete only when all of the following are true:

1. Two different provider records can belong to one canonical catalogue project.
2. Provider/external identity remains unique and traceable.
3. Field-specific source priority produces deterministic canonical values.
4. Admin overrides for approved fields survive repeated scraper/import sync.
5. Resetting an override recomputes the source-resolved value.
6. A working project cannot be created without a catalogue reference.
7. Working-project forms cannot modify catalogue-owned data.
8. Existing custom projects are backfilled without losing workflow relations.
9. Catalogue floor plans are shared by all linked working projects rather than copied per project.
10. Catalogue export from one deployment imports idempotently into another by UUID.
11. No tenant subsystem is introduced.
12. Existing brochure, campaign, owner-listing and analysis regression tests remain green.

## 13. Execution preconditions for the next Claude

The local checkout used to write this handoff is significantly behind remote branches. Re-verified 2026-07-12 (review gate):

- local `dev-chen` has **diverged** from `origin/dev-chen`: 1 local-only commit (`6840aa14` "F2f initial code" by shawn, superseded by the F2f showroom implementation already merged remotely) and 157 commits behind — a fast-forward is therefore impossible;
- `origin/dev-chen` and `origin/master` have diverged: 30 commits only on dev-chen (video/scout work), 8 only on master (including the catalogue revamp `84af8b0d`);
- `origin/master` is still at the inspected baseline `80dac964`.

Before implementation:

1. Preserve all user/untracked files; leave local `dev-chen` (and its stray commit) untouched.
2. Fetch remote state.
3. Create the feature branch from the remote head: `git switch -c feat/petahk-project-catalogue origin/dev-chen`.
4. Merge current `origin/master` into the feature branch without destructive reset/checkout.
5. Run the focused existing catalogue/project tests before changing code (confirm `phpunit.xml` still targets `petav3_testing` first).
6. Follow the linked plan in small TDD commits.

Review-gate amendments (2026-07-12) reflected in the plan:

- `catalog_floor_plans` additionally carries the derived rent statistics (`rent_low`, `rent_median`, `rent_high`, `rent_fully_furnished`, `rent_partial_furnished`) because existing consumers (`ProjectsController@renderShow`, `FloorPlan::toOption()`, the Show page) read them; `rental_price` / `market_value` remain the only override-whitelisted floor-plan fields. No `suffix` column — unit matching is workflow-owned via `OwnerListingConfig::matchRule`.
- The `projects.catalog_project_id` NOT NULL constraint is deferred to the later cleanup migration (with the compat-column drop); this release backfills all rows and enforces the invariant in `StoreRequest` and the repositories.
- The working-project `update` endpoint is kept but trimmed to the workflow-owned `highlights` field only, per §9 ("edit actions only change workflow-owned state").

Implementation plan: `docs/superpowers/plans/2026-07-12-petahk-project-catalogue.md`.

---

## Phase 2 — Project Data Completion (2026-07-13/14)

Implemented per `.claude/plans/2026-07-13-petahk-project-data-completion.md`
(rev 4, owner-approved gates G0 = option A multi-valued `market_segments`,
G9 = option A shared-GCS portable object references). Contract reference:
`docs/property-detail-contract.md`; HK column map:
`docs/hk-catalogue-mapping.md`.

### What landed

- **Contract**: `detail_contract` publication blocking lists (per segment
  membership, union for dual-segment records), four launch AI content
  locales, six-preset-to-locale routing, provider registry and staleness
  thresholds — locked by `CatalogueContractTest`.
- **Schema**: `market_segments` (set-union merged), per-country `slug`,
  `price_min/max`, `sale_status`, `completion_date`, `parking_info`,
  `management_fee`, `sale_process`, `name_translations` (per-locale keyed
  merge — one provider can never erase another's locale),
  `parent_catalog_project_id` (flat explicit grouping, validated endpoint),
  `published_at`; `catalog_floor_plans.car_parks`; new tables
  `catalog_media` (dual representation: scraped url + stable
  (source, source_key) identity, or uploaded `media_id` — signed URLs never
  persisted), `catalog_ai_contents` (append-only versions, concurrency-safe
  allocation), `catalog_sync_runs` (run history + source watermarks).
- **Packages**: SCHEMA_VERSION 2 (v1 still imports) — new fields, car
  parks, slug (create-only with deterministic remap), parent-by-uuid second
  pass, media with self-contained source identity (absent/ambiguous
  identity rejects the whole package at preflight, zero writes) and G9
  gcs_object references.
- **Ingestion**: `catalogue:sync` + `ProviderAdapter` contract +
  `CatalogueIngestionService` (savepoint per record, per-source media
  reconciliation, stale marking only after complete clean full snapshots,
  watermark staleness flags); EdgeProp adapter with the scheduled-file
  inbox producer contract; three HK adapters (house730, centanet,
  centanet-estates) over the read-only reference schemas, tested against
  redacted production payload fixtures (`tests/Fixtures/hk/`);
  `catalogue:hk-map-sources` deterministic-key mapping report (auto-attach
  behind G5).
- **AI content**: `GenerateCatalogueContent` AiJob, hash-gated with
  completion-time recheck (SUPERSEDED, never latest), triggered by
  ingestion + all admin override endpoints, behind
  `project_catalogue.ai_content_enabled` (ships false).
- **Publication + detail**: completeness validator + `catalogue:completeness`
  + publish/unpublish endpoints; `GET /{country}/projects/{slug}` shared
  detail resource (published-only public, VIEW_PROJECTS preview, locked
  analysis flags, locale routing with English fallback); admin freshness
  strip. Scheduler entries ship disabled behind
  `project_catalogue.sync_enabled` (G6).

### Open gates / not done

- **G7**: Malaysia new-project data source — EdgeProp carries no sale
  status/hero/gallery; MY new-project records are correctly blocked from
  publication until a source (or curated entry) lands (T12 pending).
- **G2**: PetaV3's `reference` connection needs read access to the
  `ih_hk_*` tables before the first real HK backfill.
- **G3 finding**: petaV2 prod schedules only `hk:warm-assets`; the HK
  scraper commands run MANUALLY today — producer ownership/cadence needs an
  owner decision (watermark staleness guards the gap meanwhile).
- **G5/G6**: HK auto-attach and sync cadence config gates ship off.
- Locale scope: catalogue content only — the hk/tw UI chrome still renders
  zh_CN until a dedicated UI translation task.
