# Catalogue transfer packages

Versioned `catalogue:export` packages, committed so another deployment can pull
one with `git pull` instead of needing an out-of-band file transfer.

These are **large binary blobs and permanent repo weight** — git keeps them
forever, even after deletion. Commit one only when it is actually being shipped,
and prefer scp/GCS for throwaway or frequently-regenerated packages.

## What is in here

| File | Rows | Raw | Exported |
|---|---|---|---|
| `catalogue-full-20260810.json.gz` | 17,541 projects | 271 MB | 2026-08-10 |

`catalogue-full-20260810.json.gz` carries the 8–10 Aug catalogue backfill:
AI/web-search `completion_year`, derived `sale_status`, floor-plan-derived
`price_min`/`price_max`, the normalised classification vocabularies, the HK
region backfill, the merged floor plans, the developers pivot, and the
`manual_overrides` pins that stop `catalogue:sync` reverting any of it.

Digests (verify after transfer):

```
gz   sha256  8c620fee16b4cbe0f5afd49683e561fa44e2657d8789d8269a88af1ec3b5fba2
json sha256  4426d9dcd7eba0df7def61132ae8f422e182b790ab61ed42d461740521016b8c
```

It was produced with `--skip-orphans`, which omitted 726 rows (690 media +
36 AI content) across 36 unpublished MY projects whose provider source row had
been deleted — ids 19969–20004, one bulk delete. Their lineage cannot be
reproduced anywhere, so they could not travel; the export prints every one.

## ⚠️ Import the PARTS, not the whole file

`json_decode` turns a package into a PHP array about **8× the file size** — the
271 MB full package peaks at **2.2 GB**. CLI `memory_limit` is normally `-1`, so
PHP never stops; the kernel OOM killer picks a victim instead, and on a web box
that is usually php-fpm or MySQL. **That is how importing a catalogue took a site
down on 2026-08-11.**

`parts/` holds the same 17,541 aggregates split 12 ways. The largest peaks at
**317 MB** instead of 2.2 GB, and each part's memory is returned to the machine
when its process exits.

```bash
gunzip -k catalogue-packages/parts/*.gz

for f in catalogue-packages/parts/*.json; do
    echo "=== $f"
    php artisan catalogue:import "$f" --resume || break
done
```

Order does not matter — no aggregate in this package has a parent, and
`linkParents` resolves parents from the database rather than the package, so a
child imported before its parent still links on a later run. Re-running any part
is safe; every write is an upsert.

**Run the loop until every part reports `0 failed` — usually twice.** Developer
identity resolves against the `developers` table as it currently stands, and the
import is building that very table as it goes: creating rows, merging aliases.
An aggregate early in the first pass can hit an ambiguity against a half-built
table and refuse, then resolve cleanly on the next pass once a later aggregate
has created the developer it needed. Observed on 2026-08-11: ~922 aggregates
failed their developer step on pass one and every one of them succeeded on pass
two, with no data changed in between. Only failures that SURVIVE a second pass
are genuine ambiguity worth investigating.

`catalogue:import` now refuses to start when a package would not fit in
available memory, and prints this loop instead. `--force-memory` overrides it.

## Importing it on the other side

```bash
git pull
gunzip -k catalogue-packages/catalogue-full-20260810.json.gz
sha256sum catalogue-packages/catalogue-full-20260810.json   # must match above

# 1. DRY RUN FIRST — read the create/update split before writing anything.
php artisan catalogue:import catalogue-packages/catalogue-full-20260810.json --dry-run
```

**Read the dry-run output before going further.** Matching is by UUID only —
there is no name/coordinate fallback:

- **mostly `would update`** → the two catalogues share identity. Proceed.
- **mostly `would create`** → the UUIDs do NOT align (that side built its rows
  from its own `catalogue:sync`). **Stop.** Importing would duplicate the
  catalogue rather than backfill it.
- **`conflict(s)`** → a provider source in the package is already attached to a
  different local canonical. The import refuses as a whole; resolve with an
  explicit admin source mapping first.

```bash
# 2. Write it (single transaction; validated before any write).
php artisan catalogue:import catalogue-packages/catalogue-full-20260810.json

# 2b. …or, when reconciling two deployments that have drifted, commit each
#     aggregate on its own and keep going past the ones that genuinely conflict:
php artisan catalogue:import catalogue-packages/catalogue-full-20260810.json --resume

# 3. Prove it landed.
php artisan catalogue:verify-sync catalogue-packages/catalogue-full-20260810.json \
    --report=storage/app/catalogue/verify-after-import.json
```

**Which mode?** The default is one all-or-nothing transaction — right for shipping
a package as a single consistent snapshot, and it means a failure leaves the
destination exactly as it was. `--resume` gives each aggregate its own
transaction and reports the failures grouped by cause at the end. Use it when the
two sides have drifted and a handful of ambiguous records would otherwise cost
the other seventeen thousand their import.

`--resume` is safe to run repeatedly. Every write is an upsert keyed on stable
identity, so a second run rewrites the same values rather than duplicating
anything — fix the reported causes, run it again, and only the remaining
failures are left.

`catalogue:verify-sync` is read-only and exits non-zero on drift, so it can gate
a deploy script. `in sync = 17,541, drifted = 0, missing = 0` means both
databases hold the same catalogue.

## Reading the verify output

**Expected divergences are contract, not drift.** The importer deliberately does
not copy everything, so the verifier does not count these as failures:

- a field the DESTINATION pinned in `manual_overrides` keeps its local value —
  a local decision outranks an incoming package;
- `slug` is workflow-owned at the destination and untouched on update;
- `manual_overrides` is merged (package ∪ local, local winning), so the
  destination may legitimately hold a superset.

Use `--strict` to count the first two as drift too, when you want byte-parity.

**Two things the verifier cannot see**, because they never travel:

1. `developer` / `developer_brand` are not in `CatalogProject::CANONICAL_FIELDS`.
   The developers PIVOT does travel and IS verified; the text columns are only a
   display fallback behind `primaryDeveloperName()`, so they stay stale on the
   destination until re-derived there.
2. `catalog_analysis_snapshots` (the precomputed unit analysis) is not part of
   the package. Re-run `catalogue:precompute-analysis --apply` on the
   destination — it is hours of engine time and costs money per run.

## Regenerating a package

```bash
php artisan catalogue:export storage/app/catalogue/catalogue-<date>.json [--country=MY]
```

⚠️ **That produces ONE file — the single large package this README's own warning tells you not
to import.** `catalogue:export` takes only `{file}`, `--country` and `--skip-orphans`; there is
**no split option**, and nothing in the repo produces the `parts/` layout described below. Until
a splitter exists, either import the whole file on a box with the headroom (`json_decode` costs
roughly 8× the file size, so a 271 MB package peaks at ~2.2 GB) or split it by country with
repeated `--country=` runs.

It aborts on rows whose lineage it cannot reproduce; add `--skip-orphans` to
omit them instead. Either way every omitted row is printed. Run
`catalogue:verify-sync` against the SOURCE database afterwards — a package must
verify clean against the database it came from, which is what proves the
comparator itself is honest.
