# AI Video Studio → "Aixcut-style" Editing Layer — Blueprint (proposal)

**Status:** Proposal — decide before building. Not an execution plan yet.
**Owner:** Chen. **Kept local** (like the other Video Studio docs, dropped from the PR in `b227838`).
**Scope (agreed with user):** 3 capabilities — *find best moments*, *edit-by-text*, *conversational agent* — plus the flow *chat → generate → edit mode*, on **both** inputs (our generated content **and** uploaded real footage).

---

## 1. TL;DR & recommendation

- The product becomes **Project-centric**: every session starts with **Create Project**. A Project = one property/listing that holds a **shared asset pool** + **many cut versions**. Assets are added *inside* the project, from two sources: **AI-generated (via chat)** or **uploaded from the computer**.
- Build in **4 phases**:
  - **Phase 0 — Project container** *(do first — it's the front door you asked for)*: `Create Project` → project workspace that holds the asset pool and all cuts. Lightweight shell; everything else lives inside it.
  - **Phase 1 — Edit mode for generated videos** *(first real feature)*: reuses the *re-package-without-Seedance* path (already proven ad-hoc when we re-fixed #20's captions), needs **no STT**, costs **~nothing per edit**. Delivers "edit by text + reorder + find best cut + instant re-render" for our own content.
  - **Phase 2 — conversational drafting + AI assets** (chat interviews the user → drafts/generates assets & storyboards into the project).
  - **Phase 3 — uploaded real-footage editor** (STT → segment → best-moments → text-based cut). The big lift ("half a video editor"); last.
- **Recommended build order: Phase 0 → Phase 1**, then re-evaluate. First demoable milestone: a Project you can create, generate a video in, then re-cut by text instantly.

---

## 2. Competitive teardown

| Tool | Core job | Signature features |
|---|---|---|
| **Aixcut** (aixcut.com) | AI video editor / "creation agent" | Upload raw footage → auto-find best moments (~60s); edit-by-text; long→short; "4 layers of AI analysis" |
| **Descript** | Text-based editor | Edit the transcript = edit the video; filler-word removal |
| **OpusClip** | Long→short repurposing | Auto-pick viral moments from a long video → vertical clips + auto-captions |
| **CapCut** | Transcript editing | STT → edit text → cuts sync to timeline |

**The three reusable mechanics behind all of them:**
1. **Edit-by-text** = STT → transcript ↔ timeline kept in sync → deleting/reordering text cuts/reorders footage.
2. **Find best moments** = STT + an LLM that scores transcript segments for hooks/highlights → auto-select.
3. **Long → short** = topic/moment detection → cut into vertical clips + auto-caption.

All three depend on one thing: **a transcript with timestamps + segment boundaries.**

---

## 3. The insight that shapes the whole plan

Our two input types are **not** equal in difficulty, because of that transcript dependency:

- **Our generated content already has the transcript for free.** Each `video_scene` carries `caption` + `voiceover` text, an exact `duration`, a `type`, and (after generation) a **stored clip** (`COLLECTION_CLIP`). So "edit-by-text", "reorder", "drop", and "find best scenes" are just **scene-level operations + a re-render from the stored clips** — **no STT, no Seedance, no extra cost.** We already proved the re-render path when we re-packaged #20.
- **Uploaded real footage has none of that.** It needs a brand-new pipeline: STT transcription with word/segment timestamps → segmentation → LLM highlight scoring → frame-accurate ffmpeg cutting → transcript↔timeline UI. This is the expensive part.

➡️ **Do the cheap, high-value half (generated content) first; defer the expensive half (uploaded footage).**

---

## 4. Target product shape & flow

```
  Create Project  (name = the property/listing)
        │
        ▼
  ┌─ Project workspace ─────────────────────────────────────────────┐
  │                                                                  │
  │  Asset pool  (shared; add anytime, inside the project)           │
  │     ● upload from computer   ● AI-generate via chat              │
  │                                                                  │
  │  Cuts  (many versions under the project)                         │
  │     Generate a cut ─▶ EDIT MODE (text-based) ─▶ Export           │
  │        (existing            • edit caption/voiceover text         │
  │         pipeline,           • reorder / drop / trim               │
  │         unchanged)          • "Find best cut" (LLM picks best N)  │
  │                             • re-render from stored clips (≈free) │
  └──────────────────────────────────────────────────────────────────┘
```

Unifying ideas:
- **Project = one property** → a shared **asset pool** + many **cut versions**. (Supersedes today's flat generation list + `root_id` version family.)
- **Assets are project-level and reusable** across cuts; added *inside* the project, either uploaded or AI-generated via chat. (Today media is bound to a single generation — this lifts it to the project.)
- **Everything renderable is a list of editable "segments."** A generated scene is a segment; an uploaded-footage highlight is a segment. The editor edits segments; the renderer stitches them. Phase 3 just produces segments from a new source (uploaded video + STT) and feeds the *same* editor + renderer.

---

## 5. Phased roadmap

### Phase 0 — Project container *(do first — the "Create Project" front door)*

**Goal:** every session starts at a project list with a **Create Project** button; a project is the workspace that owns a shared **asset pool** and all **cut versions**.

- **Capabilities delivered:** the project-centric structure + entry point. No Aixcut editing features yet — this is the foundation everything else hangs off.
- **Data model:**
  - New `projects` table — `id` + `uuid` (HasUuid), `name` (the property), `brief` (JSON: chat-gathered requirements, filled later in Phase 2), `status`, blame + timestamps + soft-deletes (standard key model per GUIDELINES).
  - `video_generations.project_id` — a **cut belongs to a project**. The existing `root_id`/`revised_from_id` version family now lives *within* a project.
  - **Shared asset pool:** lift input media from generation-bound to **project-bound** (attach media to the `Project`; a cut references which pool assets it uses). Each asset records its **source** (`uploaded` | `ai_generated`). *This media re-ownership is the meatier part — can be staged: 0a = project shell + cuts belong to a project; 0b = shared asset pool.*
  - Migration: backfill existing generations into a per-family (or single "legacy") project so nothing is orphaned.
- **Backend:** Project CRUD following the `manage` module pattern (`routes/web.php` group + `Manage\Video\ProjectsController` + repository in `Src\Video`); the existing video list/flow becomes project-scoped.
- **Frontend:** a **Projects** index (list + Create Project modal) and a **Project detail** workspace (tabs: *Assets* | *Cuts*). Today's `Manage/Video/Index` flow moves inside the project detail.
- **Reuses:** the shared DataTable/index pattern (GUIDELINES §14), MediaService, existing video flow (re-parented under a project).
- **Effort:** **Medium** (mostly the media re-ownership + migration). **Risk:** Low–Medium. **Cost:** ~0.
- **Verify:** create a project → add an asset (upload) → generate a cut *inside* it → the cut and asset both belong to the project; the project lists its cuts + asset pool; existing videos still reachable under their backfilled project.

### Phase 1 — Edit mode for generated videos *(first real feature, after Phase 0)*

**Goal:** after a video is READY, the user re-enters an editor on the *already-generated* clips and re-cuts by text, instantly and for free.

- **Capabilities delivered:** edit-by-text (the scene caption/voiceover *is* the transcript), reorder, drop/trim scenes, **"Find best cut"** (Gemini scores the scenes and proposes a keep-set/order for a target length, e.g. best 4 of 6 → a 20s version).
- **Render = re-package, no Seedance:** reuse stored `COLLECTION_CLIP` per scene → `ClipNormalizer` → `FfmpegStitcher` → `CaptionRenderer` + `VoiceoverService` + `VideoPackager`. (Exactly the path we ran by hand for #20.) Dropped scenes are skipped; reordered scenes restitch; text changes re-render captions/VO only.
- **Data model:** none new strictly required — `video_scenes` already holds text/order/duration/type and clips are stored. Add a soft flag for "scene included in cut" (or just delete the scene row on drop, since clips persist by position). Optionally store a `cut_label`/version note.
- **Backend:** `POST /manage/video/{id}/repackage` — a `RepackageVideoJob` (or `GenerateVideoJob` in a "reuse clips" mode) that skips Seedance and rebuilds from stored clips. `POST /manage/video/{id}/best-cut` — Gemini scores scenes → returns suggested keep/order (no mutation; the UI applies it).
- **Frontend:** an **Edit** action on a READY card → opens the storyboard editor in "post" mode: clips shown, text editable, drag-reorder, drop, "Find best cut" button, "Re-render (no AI cost)" button. Largely reuses `StoryboardStudio`/`SceneCard`.
- **Reuses:** re-package path, scene model, storyboard editor, Gemini (`AiClient`), Horizon.
- **Effort:** **Small–Medium.** **Cost/edit:** ~zero (no Seedance). **Risk:** Low.
- **Verify:** edit text + drop a scene + reorder → re-render → output reflects changes, length correct, no Seedance call in logs. `find-best` returns a sane keep-set.

### Phase 2 — Conversational drafting (chat → storyboard)

**Goal:** replace/augment the upload-form with a short AI conversation that gathers the brief and then drafts the storyboard.

- **Capabilities delivered:** the "conversational agent" front-end of the flow.
- **Backend:** a chat endpoint backed by Gemini (`AiClient`) that asks for property type / highlights / tone / length / language / presenter, then calls `StoryboardDraftService::draft(...)` to produce scenes. Optionally persist the chat transcript on the generation for later context.
- **Frontend:** a chat panel; on "looks good" → draft → drop straight into Phase 1's edit mode.
- **Reuses:** `StoryboardDraftService`, Gemini.
- **Effort:** **Medium.** **Risk:** Low–Medium (mostly prompt/agent design + making the chat reliably terminate in a valid draft).
- **Verify:** a scripted conversation reliably yields a valid N-scene draft with the chosen settings.

### Phase 3 — Uploaded real-footage editor *(the big lift — defer)*

**Goal:** upload real video → transcribe → find best moments / long→short → edit-by-text → assemble.

- **New pipeline:** STT (word/segment **timestamps** are mandatory) → segment (sentence/pause/scene) → LLM highlight scoring → frame-accurate ffmpeg cutting → transcript↔timeline editor.
- **Data model:** a new **source asset** + **segments** table (`start`, `end`, `text`, `score`, `kept`). Kept segments map onto the existing CLIP-scene/segment model so they flow through the *same* editor (Phase 1) and renderer.
- **Backend:** upload → `TranscribeJob` → segments; `best-moments` scoring; `apply edits` → cut → assemble (reuse stitch/package). Long→short = auto-pick top-scored segments into a vertical short.
- **Frontend:** transcript editor synced to a player + timeline (the hard UX); highlight suggestions; export.
- **Effort:** **Large** (this is a genuine video editor). **Risk:** High — STT timestamp accuracy, transcript↔timeline sync UX, ffmpeg frame-accuracy, per-minute STT cost, long-video performance.
- **Verify:** upload a 2-min clip → transcript with timestamps → delete sentences → cut output matches; "find best" produces a coherent 30s short.

---

## 6. Risks & decisions to make

| Topic | Options / risk | Recommendation |
|---|---|---|
| **STT** (Phase 3) | Gemini audio transcription (already integrated) vs Whisper (self-host) vs Deepgram (paid, great timestamps) | Prototype with **Gemini** first; switch to Deepgram/Whisper if word-timestamps are weak |
| **Best-moments scoring** | Gemini over transcript text; optionally add visual cues | Start **text-only Gemini**; visual later |
| **Transcript ↔ timeline sync** | Needs reliable per-word/segment timestamps; hardest UX | Prototype this **first** in Phase 3 before building the editor |
| **ffmpeg cutting** | Frame accuracy vs re-encode cost | Segment-level cuts; reuse existing helpers |
| **Cost** | Phase 1 ≈ 0 (reuse clips); Phase 3 = STT per minute + storage | Phase accordingly; gate uploads by length |
| **Scope creep** | Phase 3 is 5–10× Phase 1 | Don't start Phase 3 until Phase 1 proves the edit UX + demand |

---

## 7. Recommendation & decision asks

**Recommendation:** **Phase 0 → Phase 1.** Phase 0 stands up the project-centric structure the user wants (Create Project + workspace + shared asset pool); Phase 1 then delivers the first real Aixcut-style win (edit-by-text + find-best-cut + instant, free re-render) for our generated cuts, reusing the no-Seedance re-package we already have. Phase 2 (chat + AI assets) next; Phase 3 (uploaded-footage editor) is a separate, larger greenlight afterwards.

First demoable milestone: **create a project → add a photo → generate a cut → re-cut it by text in seconds (no AI cost)**.

**Decide:**
1. Confirm build order **Phase 0 → Phase 1**? (If yes, next step is a detailed Phase-0 execution plan + TDD.)
2. Phase 0 asset pool: do it **all at once** (0a shell + 0b shared pool together) or **stage** it (shell first, shared pool right after)?
3. Phase 3 timing: is "uploaded real footage" needed soon, or is "edit our generated videos under projects" (Phase 0–2) enough for the near-term goal? (Plus STT provider preference for when we do it: Gemini vs Deepgram/Whisper.)
