# AI Video Studio — Project Chat-to-Storyboard (v1) — Frozen Spec & Plan

**Status:** Spec **frozen**, pre-implementation. Branch `dev-chen`. Builds on Phase 0 (Projects container) and C-3 (conversational drafting). Kept local like the other AI Video docs.
**Date:** 2026-06-24. **Owner:** Chen. **Reviewer / merge:** Lee Jie.
**Read first:** [`ai-video-phases-abcd-handoff.md`](ai-video-phases-abcd-handoff.md) (C-3 today), [`ai-video-aixcut-blueprint.md`](ai-video-aixcut-blueprint.md).

Reshapes today's "AI Draft" tab into a **chat-first guided flow**: an empty project opens straight into a conversation that interviews the salesperson, collects photos, proposes a readable storyboard (方案), lets them refine it in plain language, and on confirmation drops the finalized scenes into the existing StoryboardStudio editor.

---

## 1. Product definition (frozen)

1. **Entry — empty project = full-page chat.** Opening a project with **no cuts** lands directly in the AI chat (not the tabbed table). Once the project has cuts, the page shows the normal Cuts/Assets tabs and the chat becomes a secondary entry (a "Chat with AI" button/tab).
2. **Guided interview.** The AI plays a real-estate ad director, asks **one focused question at a time**, and **actively asks the customer to send photos** (videos welcome too).
3. **Attachments in chat → project pool.** Photos/videos are uploaded from the chat composer (attach button), land in the **project asset pool** (persisted), and render inline in the transcript.
4. **AI understands the photos.** Each uploaded photo is summarized once (vision), so the proposed storyboard maps scenes to **specific photos** ("第2镜：你这张客厅落地窗图").
5. **No photo → ask for one.** If the customer hasn't sent any photo, the AI **requests one in conversation**. No auto image-generation, no hard block — it just keeps asking. Finalizing needs ≥1 photo.
6. **Readable 方案 = the storyboard.** When the AI has enough (info + ≥1 photo), it proposes a **readable, scene-by-scene 方案** that is, underneath, the structured storyboard.
7. **Refine in plain language.** The customer adjusts by talking ("把第二镜改成强调学区", "加一个泳池镜头"); the AI re-emits a revised 方案. Loop until happy.
8. **"就这么定" = one-way handoff.** Confirming persists the finalized scenes into a **new draft generation** (referenced pool photos copied into the generation's input media) and **redirects to StoryboardStudio** (the pre-generation editor). After this, **the chat no longer edits that draft** — all further editing happens in StoryboardStudio. Starting over = a fresh conversation producing a fresh draft.
9. **Videos (v1).** Sent videos go to the pool and the AI acknowledges them, but the auto-drafted 方案 is **photo-driven** (presenter scenes). Videos are added as CLIP scenes **manually in StoryboardStudio**.

---

## 2. v1 technical boundaries (frozen)

- **Image understanding = summarize-on-upload, not per-turn base64.** When a photo is uploaded, run **one** Gemini vision call → a short structured summary, cached on the media `meta` (`meta.summary`). Interview turns send **text summaries only** (cheap). The final scene-drafting step **may** re-feed the raw images only if summary-driven quality proves insufficient (held as an enhancement; v1 default is summary-driven).
- **Confirm is a one-way handoff.** Pre-confirm the **chat owns** the transient 方案 (front-end state); post-confirm the **StoryboardStudio owns** the draft generation. The chat never re-opens or re-edits a committed draft → single writer, no split state.
- **Transcript ephemeral.** No messages table in v1. Reloading mid-chat resets the conversation; **uploaded photos persist** in the pool (so a reload doesn't lose materials).
- **Scene count dynamic** (~3–6), decided by the AI from the number of photos/selling points — not the hardcoded 4.
- **Defaults preserved:** `zh` / `zh_female_warm` / presenter on / `9:16` / current resolution. All editable later in StoryboardStudio. The interview's gathered *tone* feeds scene prompts, not the voice pick.
- **The new chat replaces the old "AI Draft" tab** and its manual "Draft storyboard from N pool images" button.
- **Stateless server** (matches today's interview): the front-end holds the running transcript + latest 方案 and re-sends them; the server persists only on confirm.

**Out of scope for v1:** video auto-orchestration into the 方案, Seedream auto-fallback when no photo, transcript persistence, editing a committed draft from chat, multi-photo-per-scene composition.

---

## 3. Flow & endpoints

```
Open project (no cuts) ─► full-page chat
        │
        ▼
[chat turn]  POST projects/{id}/chat        { messages }                → { reply }            (a question)
[attach]     POST projects/{id}/chat/upload { file }       → store to pool + 1 vision summary  → { asset: {uuid,url,kind,summary} }   (JSON, inline in transcript)
        │  (controller injects current pool photos' summaries into the interview each turn)
        ▼
[ready]      POST projects/{id}/chat        { messages }                → { ready:true, scenes:[…], brief }   (the 方案)
[refine]     POST projects/{id}/chat        { messages(+feedback) }     → { ready:true, scenes:[…] }          (revised 方案)
        │
        ▼
["就这么定"] POST projects/{id}/storyboard/commit { scenes, … }        → create draft gen, copy referenced pool photos to video_input,
                                                                          persist scenes, redirect → /manage/video?draft=<uuid>&project=<uuid>
        ▼
StoryboardStudio (existing) — edit → Generate.  Chat does not touch this draft again.
```

- **Pool image ordering is stable** (by `id`) for both the interview injection and the commit mapping, so a scene's `image_index` resolves to the same photo in both places.
- `chat()`'s `ready` response can recur across turns (refinement) — any scenes-bearing response is "here is the (updated) 方案"; the front-end keeps the input open for more refinement and shows the **就这么定** button.

---

## 4. File-by-file plan

### Backend

1. **`src/Video/Services/PropertyInterviewService.php`** (rewrite the contract)
   - `next(array $messages, array $poolImages): array` — `$poolImages` = `[{index, summary, kind}]` injected by the controller.
   - System instruction: interview one Q at a time; **explicitly ask for photos**; once enough (+≥1 photo summary present), output the **structured storyboard JSON** (`{ready:true, brief, scenes:[{image_index, prompt(EN), caption(CN), voiceover(CN)}]}`) — same scene shape as `StoryboardDraftService`. On subsequent turns with feedback, re-emit revised scenes.
   - Keep the tolerant `{ready:…}` JSON parsing.

2. **`src/Video/Services/StoryboardDraftService.php`** (extract shared shaping)
   - Pull the **scene-row shaping** (presenter-prepend via `VideoTemplate::presenter`, `duration_seconds`, `image_index → media id` mapping, fallback) into a small reusable method/service so **commit** persists scenes through the *same* shaping path the drafter uses. No behavior change to the existing `draft()`.

3. **`src/Video/Services/ImageSummaryService.php`** (new, small)
   - `summarize(string $absolutePath): string` — one `GeminiClient::imagePart()` + flash call → a short structured CN summary ("客厅·落地窗·城市景观·自然光·现代风"). Defensive: returns `''` on failure (upload still succeeds; AI just lacks that summary).

4. **`app/Http/Controllers/Manage/Video/ProjectsController.php`**
   - `chat()` — inject the project's pool-photo summaries (ordered by `id`) into `PropertyInterviewService::next()`; return its result as JSON. (No more persisting `brief` here is fine, or keep persisting brief on ready.)
   - `chatUpload()` (new) — validate one image/video, store to pool (`COLLECTION_ASSET`, `meta.source='chat'`); for images run `ImageSummaryService` and store `meta.summary`; return the asset as **JSON** (uuid, url, kind, summary) for inline rendering.
   - `commitStoryboard()` (new) — `ownedProject`; validate posted `scenes` (+ ratio/resolution/voice/language/template defaults); create a **draft** `VideoGeneration` (`project_id`); copy each referenced **pool photo** into `video_input` (reuse the copy logic from `VideoGenerationsController::draft`); persist scenes via the shared shaping + `VideoSceneRepository::createMany`; redirect to `manage.video.generations.index` with `?draft=<uuid>&project=<uuid>`.
   - `show()` — add `has_cuts` (or reuse `cuts` count) so the page can decide chat-first vs tabbed.

5. **Form Requests** (new, under `app/Http/Requests/Manage/Video/Projects/`)
   - `ChatUploadRequest` (single `file`: image/video, size cap), `CommitStoryboardRequest` (`scenes` array shape + bounds, defaults). Reuse the 9-image cap server-side.

6. **`routes/web.php`** — add `projects/{id}/chat/upload` (`projects.chat.upload`), `projects/{id}/storyboard/commit` (`projects.storyboard.commit`). Keep existing `projects.chat`.

### Frontend

7. **`resources/js/Pages/Manage/Video/Projects/Show.vue`** (main UI work)
   - **No cuts → render the full-page chat** (no tabs). **Has cuts → tabs (Cuts/Assets) + a "Chat with AI" entry** that opens the same chat.
   - Chat composer gets an **attach button** → `chat/upload` (XHR) → push an image/video bubble into the transcript and into the local pool list.
   - When a turn returns `scenes`, render the **readable 方案** (per scene: photo thumbnail · 第N镜 · 字幕「caption」· 旁白「voiceover」), keep the input open for **plain-language refine**, and show **就这么定**.
   - **就这么定** → POST `storyboard/commit` with the latest `scenes` (Inertia visit → server redirects to StoryboardStudio).
   - Remove the old `draftFromChat()` / "Draft storyboard from N pool images" path.

### Tests

8. **Unit**
   - `PropertyInterviewServiceTest` — asks-a-question vs emits-scenes (stub Gemini); requires a photo summary before going ready; refine re-emits scenes.
   - `ImageSummaryServiceTest` — returns summary on success, `''` on failure.
   - shaping reuse test (presenter prepend + `image_index → media id`) stays green.
9. **Feature**
   - `ProjectsControllerTest` — `chatUpload` stores to pool + writes `meta.summary` (stub vision) + returns JSON; `commitStoryboard` creates a draft gen, copies referenced pool photos to `video_input`, persists scenes, redirects with `?draft=&project=`; **owner scoping** (foreign project/asset → 404); rejects commit with 0 photos.
10. **Manual e2e (Chen)** — empty project → chat → ask-for-photo → upload → 方案 → refine → 就这么定 → lands in StoryboardStudio → Generate.
11. **Smoke** — one real chat turn + one real upload summary against ARK/Gemini (token-cost sanity for the summarize-on-upload path).

---

## 5. Verification gates

- `"$HOME/Library/Application Support/Herd/bin/php" -d memory_limit=512M vendor/bin/phpunit tests/Unit/Video tests/Feature/Video` green (test DB `petav3_testing`, RefreshDatabase — **never** `migrate:fresh` the dev DB).
- `npm run build` clean.
- **Horizon**: no new always-on job in v1 (summary is synchronous on upload). If summary is moved to a job later, `horizon:terminate` + restart after touching it.

---

## 6. Open implementation notes (decided, not blocking)

- Summary call is **synchronous on upload** (single image, like `generateAsset`). If it adds noticeable latency, move to a queued job later (deferred).
- `commitStoryboard` reuses `VideoGenerationsController::draft`'s pool-photo copy + scene-shaping; if that means sharing code across the two controllers, extract the copy/shaping into the Video service layer rather than duplicating.
- Readable 方案 is rendered **client-side from the structured scenes** (caption/voiceover/photo) — the AI is not asked for extra prose, keeping the output schema single-source.

---

## 7. Addendum — chat uploads: batch, PDF, staged-then-send (2026-06-24)

Refinements to the in-chat upload after the initial build:

- **Batch upload.** The composer accepts **multiple files at once**. `ChatUploadRequest` is now `files[]` (≤12, each ≤100MB); `chatUpload()` loops and returns `{ assets: [...] }` (was a single `{ asset }`).
- **PDF → extracted photos.** A dropped **brochure PDF** runs through the existing `BrochureExtractionService` (`ingestPdf()` — the same proven path as `VideoGenerationsController::draft`), extracting up to 9 room photos into the pool, each summarized, each returned as an image asset. Videos store as-is; images summarize on upload.
- **Staged-then-send (no auto-send).** Picked files **wait in the composer** as removable thumbnails; nothing uploads until the user **sends** (types a message, or just presses Enter). `send()` then uploads the staged batch, pushes a bubble per stored asset, and runs one chat turn with the typed text — or a short synthetic note (`(我发了 N 个素材…)`) when the user attached without typing. The earlier "auto-upload + auto-acknowledge on pick" behavior is removed.

Tests: `test_chat_upload_accepts_multiple_images_at_once`, `test_chat_upload_extracts_pool_photos_from_a_pdf` added; the single-file upload tests updated to `files[]`/`assets[]`. Video suite **190 green**, build clean. PDF extraction + per-image summary run **synchronously** on send (bounded by the 9-image cap) — if a big brochure makes send sluggish, move extraction/summary to a queued job later.
