# AI Video — Custom Presenter Library + Voice/Caption Toggles — Frozen Spec & Plan

**Status:** ✅ **Implemented** (TDD) on `dev-chen`, not yet committed. Video suite **212 green**; full suite 630 with the same 7 pre-existing unrelated failures; `npm run build` clean. Frontend e2e + commit pending Chen. Builds on the chat-to-storyboard work ([`ai-video-chat-draft-plan.md`](ai-video-chat-draft-plan.md)) and the existing presenter/template + scene-driven pipeline.
**Date:** 2026-06-24. **Owner:** Chen. **Reviewer / merge:** Lee Jie.

Two requested additions: (A) per-video **voice & caption on/off toggles**, and (B) a per-user **custom AI presenter library** so every video that user generates keeps the same on-screen person.

---

## 0. The hard constraint (why B works the way it does)

Seedance/BytePlus **blocks reference images containing a real human face** (privacy review) — confirmed in `config/video.php` and `reference_seedance_byteplus_gotchas`. Person consistency today is achieved by prepending a **fixed text persona description** to every GENERATED scene prompt. So a "custom presenter from a photo" is built as: **upload photo → Gemini vision writes a locked person description → that description (not the photo) drives every video.** The photo is only the description source + a UI avatar; it is never sent to Seedance.

---

## 1. Frozen decisions

### A — Voice & caption toggles
- **Two whole-video switches**: 要配音 (`voiceover_enabled`) and 要字幕 (`captions_enabled`). Not per-scene.
- Live in **StoryboardStudio** (the pre-generation editor); **default both ON**. The chat-commit path defaults both on; the user flips them in the editor before Generate.
- Stored on the generation `options`. The pipeline **skips TTS** (silent narration) when voiceover is off, and **skips caption overlay** when captions are off. **Per-scene `caption`/`voiceover` text is kept** either way, so toggling back on just works.

### B — Custom presenter library
- A per-user library: a user can make **several** presenters and **pre-select one as active**; generation uses the active one automatically (no per-video picker in v1). No active presenter → fall back to the global `config('video.presenter')` / template (today's behavior).
- **Create** = upload **one photo** + a name → Gemini vision generates a **locked person description** (appearance / age / clothing / vibe, in the canonical persona style) → **editable** → save. The photo becomes the presenter's **avatar** (display only).
- Managed on a dedicated **"AI Presenters" page** under the Video area: grid of presenters (avatar + name + "active" badge), create / edit (name + description) / delete / set-active.
- The active presenter's **description** replaces the template persona when shaping scenes, so it is prepended to every scene → consistent person across all that user's videos.

---

## 2. Backend plan

### A — toggles
1. **`ProjectsController::commitStoryboard` + `VideoGenerationsController::draft`** — seed `options.voiceover_enabled = true`, `options.captions_enabled = true` on the draft.
2. **`StoryboardController@updateStoryboard`** (or the scene-save endpoint) — accept + persist the two booleans onto `options` (validated in the Storyboard `UpdateRequest`).
3. **`GenerateVideoJob` + `RepackageVideoJob`** — read the flags from `$gen->options`: when `voiceover_enabled` is false, **skip `VoiceoverService`** and pass `voiceoverPath = null` to `VideoPackager::package`; when `captions_enabled` is false, pass an **empty `$captions`** array (no overlay). Default missing flag = true (back-compat for existing cuts).

### B — presenter library
4. **Migration `video_presenters`**: `bigIncrements id`, `uuid`, `created_by` (index), `name`, `description` (text), `is_active` (bool, default false), timestamps, `softDeletes`. Model uses `HasUuid` + `RecordsBlame`. Avatar photo stored via **MediaService** on the presenter (`COLLECTION_AVATAR = 'presenter_avatar'`).
5. **Model `Src\Video\VideoPresenter`** (extends `Diver\Database\Eloquent\SoftDeleteModel`, `HasUuid`, `RecordsBlame`): `media()` morph relation; static `activeFor(int $userId): ?self`.
6. **`VideoPresenterRepository`** (+ Facade): `create`, `update`, `delete`, `activate` — all `DB::transaction`. `activate` clears `is_active` on the user's other presenters then sets it on this one (single-active invariant).
7. **`PresenterDescriptionService::describe(string $absolutePath): string`** — Gemini vision (reuses `GeminiClient::imagePart`) → a locked person description in the `config('video.presenter')` style; returns `''` on failure (the create flow surfaces an error instead of 500).
8. **`Manage/Video/PresentersController`** (owner-scoped by `created_by`): `index` (Inertia page), `store` (photo → `PresenterDescriptionService` → create + store avatar), `update` (name/description), `destroy`, `activate`. Mirrors `ProjectsController`'s ownership pattern (`ownedPresenter` 404s on foreign).
9. **Form Requests** `Manage/Video/Presenters/{Store,Update}Request`: store = `photo` (image, ≤100MB) + `name`; update = `name` + `description`.
10. **Routes** `routes/web.php`: `video/presenters` index/store, `video/presenters/{id}` update/destroy, `video/presenters/{id}/activate`.
11. **Integration with shaping** — `StoryboardDraftService::shape(..., ?string $presenterOverride = null)`: use the override persona when provided, else `VideoTemplate::presenter($template)` (unchanged default). `draft()` + `commitStoryboard()` resolve `VideoPresenter::activeFor(auth()->id())?->description` and pass it as the override, so the active presenter wins; no active presenter → today's template behavior.

---

## 3. Frontend plan

12. **`Pages/Manage/Video/Presenters/Index.vue`** — the library: grid of presenter cards (avatar, name, "Active" badge, set-active / edit / delete), a **Create** modal (upload one photo → preview → "Generate description" → editable textarea + name → Save). Reuses `Modal` / `ConfirmModal`. Add a nav entry "AI Presenters" in the Video area.
13. **StoryboardStudio** — two switches (要配音 / 要字幕) bound to the generation's `options.voiceover_enabled` / `captions_enabled`, default on, persisted via the existing storyboard-save. No other editor change.

---

## 4. Tests (TDD, RED→GREEN)

- **`PresenterDescriptionServiceTest`** — returns the description on success; `''` when the image is unreadable / Gemini fails; sends the image to Gemini.
- **`VideoPresenterRepositoryTest`** — create; `activate` enforces single-active (others cleared).
- **`PresentersControllerTest`** — `store` creates a presenter + avatar + description (stubbed vision), owner-scoped; `activate` flips active; `destroy`; foreign project/presenter → 404.
- **`StoryboardDraftServiceTest`** — `shape()` uses `presenterOverride` when given, else the template.
- **commit/draft** — when the user has an active presenter, the shaped scene prompt begins with the **presenter's** description, not the global default.
- **`GenerateVideoJob` / `RepackageVideoJob`** — `voiceover_enabled=false` → no TTS, `voiceoverPath` null; `captions_enabled=false` → empty captions passed to the packager; both default true when absent.

## 5. Verification gates
- `vendor/bin/phpunit tests/Unit/Video tests/Feature/Video` green (DB `petav3_testing`, RefreshDatabase — never `migrate:fresh` the dev DB).
- `npm run build` clean. Frontend e2e by Chen.
- Horizon: presenter-description vision runs **synchronously** on presenter create (one image, like the chat summary) — no new always-on job. Job-class edits (`GenerateVideoJob`/`RepackageVideoJob`) → `horizon:terminate` + restart.

## 6. Sequencing
Toggles (A, small, self-contained) first; then the presenter library (B, the bulk). Both on `dev-chen`.

## 7. Out of scope (v1)
Per-scene toggles; multiple photos / multi-angle presenter; binding a default voice to a presenter; per-video presenter override (the active selection is standing); sending any photo to Seedance.
