From d2cd4bacd342680343c19e6c9c62b3488ef4b5a0 Mon Sep 17 00:00:00 2001 From: Natasha Ann Lum Date: Fri, 10 Jul 2026 14:47:25 +0800 Subject: [PATCH 1/7] docs: scaffold product discovery folder and draft Problem Hypothesis + Lean Canvas Adds a portable product/ directory (Problem Hypothesis, Lean Canvas, Stakeholder Map, Customer Interviews, Journey Map, Personas, Opportunities templates) meant to be carried into the new repo alongside the technical RFC. Problem Hypothesis and Lean Canvas are drafted with real detail: the B2B2B subcontracted-editing partnership in negotiation, existing-alternative pricing benchmarks, and the local-AI cost/sustainability angle as a structural unfair advantage. Both are first-pass hypotheses pending cofounder review and customer interviews, not validated conclusions. Co-Authored-By: Claude Sonnet 5 --- product/00-problem-hypothesis.md | 49 ++++++++ product/01-lean-canvas.md | 106 ++++++++++++++++++ product/02-stakeholder-map.md | 35 ++++++ product/03-customer-interviews/README.md | 39 +++++++ product/03-customer-interviews/synthesis.md | 37 ++++++ .../transcripts/README.md | 22 ++++ product/04-journey-map.md | 38 +++++++ product/05-personas.md | 33 ++++++ product/06-opportunities/README.md | 31 +++++ product/README.md | 30 +++++ 10 files changed, 420 insertions(+) create mode 100644 product/00-problem-hypothesis.md create mode 100644 product/01-lean-canvas.md create mode 100644 product/02-stakeholder-map.md create mode 100644 product/03-customer-interviews/README.md create mode 100644 product/03-customer-interviews/synthesis.md create mode 100644 product/03-customer-interviews/transcripts/README.md create mode 100644 product/04-journey-map.md create mode 100644 product/05-personas.md create mode 100644 product/06-opportunities/README.md create mode 100644 product/README.md diff --git a/product/00-problem-hypothesis.md b/product/00-problem-hypothesis.md new file mode 100644 index 0000000..b7848b5 --- /dev/null +++ b/product/00-problem-hypothesis.md @@ -0,0 +1,49 @@ +# Problem Hypothesis + +Write this yourself, fast, before doing any research below. This is your working claim — not yet validated, meant to be challenged by what follows (the canvas, the interviews). Timebox: 30–60 minutes. If you're stuck on a prompt, write "I don't know yet, but I believe X" and move on. + +## What's broken today? + +*From your own experience building and running ragTech's pipeline — be specific. Not "editing podcasts is hard" but the actual moments that cost you time or made you want to quit. (You already have a technical version of this in `docs/rfcs/0001-native-desktop-rewrite.md` — this should be the human/time-cost version, not the architecture version.)* + +- Editing long-form podcasts is a timely process with many repeated elements, especially so for podcasters who are not yet established and have full-time work commitments. +- We are often using free softare like Capcut or Microsoft Clipchamp due to lack of funds, which can be buggy and take up more editing time +- It takes a dedicated half a day at least to do a standard 30min-1H podcast edit, consindering we have to clean audio, sync audio and video, cut between speakers, add intro and name title cards, and outro +- Creating short-form content through clipping for publicity from the long-form takes another day, especially when you want to create enough to distribute and promote across 1-2 weeks +- our ragTech team rotates the task amongst three of us every two weeks per episode, but even then the process is tedious. We want to focus on the content, on engaging with our audience and our products, not just social media marketing and editing. +- We don't have the funds to outsource this, and we also want to retain consistency and quality across edits +- Three of us have different editing styles and tastes, but we want to be consistent throughout and hold ourselves to high standards. However those standards inevitably falter when we don't have time. Before we built the internal prototype, each of us edited with our own software/taste and consistency suffered noticeably; the prototype has already improved this somewhat. Consistency is a problem we want the product to actively solve (e.g. house style/brand rules baked in so output quality doesn't depend on who's editing that cycle) — but it's a lower priority than raw editing speed. +- We have an intermediary prototype that works for our own videos, but it is hacky. Now we want to scale up to editing for other podcasts so this can be a viable business + +## Who else likely has this problem? + +*Hypothesis, not yet validated. Solo creators? Small teams like yours (2-3 person podcasts)? Agencies editing for multiple clients? Be as specific as you can about who — "podcasters" is too broad to interview against later.* + +- Small podcast teams just starting out — small and growing, not yet established. (400-500 followers is a rough observation from podcasters we personally know, not a strict threshold; the defining trait is "small and growing," not a specific follower count.) Split further into: (1) founders who have side businesses and income, and pay podcast studios to edit their videos and create short-form. (2) people with full-time work trying to start a podcast, editing on their own with their own resources and software like Riverside. +- Companies who want to start their own podcasts for marketing and lead generation. +- Both of the above are live/plausible segments for us right now — we haven't picked one over the other, and which we chase first may partly depend on who approaches the podcast studio rental partnership below rather than a deliberate choice we've made ourselves. +- Podcast studio rental companies who want to offer editing services to their own clients (we were reached out by one and are in the midst of securing the partnership). The shape of this partnership: they route editing work to us as a subcontractor — we do not expect them to operate our tool themselves; we (ragTech) use it internally to fulfill the work. + +## Why now? + +*What's changed that makes this worth building now rather than 2 years ago, or worth someone switching tools for now?* + +- Content creation has become more accessible, including podcast creation, due to social media and editing tools +- More companies are using podcasts for marketing and lead generation +- podcast rental studios are becoming more popular and automated (https://www.thelfgpod.com/), as well as affordable, but editing services for long-form are still largely manual and require expertise. Poddster for example, charges ~$200+ for a full episode edit, and another $200+ for 3 shortform clips from there. +- AI is being used for video editing increasingly, but commercial/cloud AI models are expensive per-token and carry real ethical/sustainability costs. ragTech is branded around being responsible technologists — building on local, efficient AI models (transcription, diarization, etc. running on-device rather than through cloud APIs) gives us a legitimate, concrete cost and sustainability advantage over AI-heavy competitors, not just a philosophical stance. + +## What does success look like in 12 months? + +*Concrete, not aspirational. E.g. "ragTech renders episodes in under 1 hour instead of 8" and/or "10 other podcast teams are using this weekly."* + +- ragTech renders episodes in under 1 hour instead of 8 +- ragTech operates as a **service business powered by internal tooling** — we are editing for 10 other podcast teams weekly, delivered as subcontracted/outsourced editing work (e.g. via the studio-rental partnership), not as a self-serve product those teams operate themselves. +- Beyond 12 months (not a year-1 commitment): if the service model validates real demand and pricing, we want the option to evolve toward a self-serve product other teams operate directly. The service business is how we keep improving the product with real usage while we figure out if/when that transition makes sense. + +## Explicit non-goals + +*What are you deliberately NOT trying to solve, at least for v1? (E.g. not a general-purpose NLE, not competing with Premiere/Resolve, not supporting every camera format.) This matters more than it seems — scope creep here becomes scope creep in the engineering RFC.* + +- Not solving for generic video-editing — scoped to podcast-style content: single-host/interview formats and multi-camera/multi-speaker conversational shows. In scope: both the long-form episode edit AND repurposed derivatives from that same source footage (short-form clips, marketing cuts) — this is not long-form-only, it explicitly includes content repurposing/marketing use cases. +- Not GA'ing the *software* to self-serve external users this year — the tool stays internal, operated only by our team. What we sell externally in year one is the editing *service* (subcontracted work, e.g. via the studio-rental partnership), not the app itself. A self-serve product other teams operate directly is a possible future direction (see 12-month success metric above), not a v1 goal. diff --git a/product/01-lean-canvas.md b/product/01-lean-canvas.md new file mode 100644 index 0000000..a0c7c29 --- /dev/null +++ b/product/01-lean-canvas.md @@ -0,0 +1,106 @@ +# Lean Canvas + +Fill this out in one sitting, ~45 minutes, all three cofounders in the room if possible. Don't overthink any single box — the point is to get your current best guess down so it can be tested, not to be right on the first pass. Expect to revise this after the first few customer interviews; that revision is the artifact working as intended, not a failure. + +Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 boxes, fill in roughly this order (1 → 9): + +## 1. Problem + +*Top 1-3 problems. Pull directly from `00-problem-hypothesis.md`.* + +- Editing a standard 30min-1hr multi-cam podcast episode takes a full half-day minimum (audio cleanup, sync, cutting between speakers, intro/outro), plus another full day to cut short-form clips for 1-2 weeks of promotion — unsustainable for teams with full-time jobs or limited budget. +- Affordable DIY tools (CapCut, Clipchamp) are buggy and slow editing down further; professional alternatives are priced per-episode in a way that doesn't scale with volume, and hiring in-house/freelance editors isn't affordable for small/early teams. +- Editing quality/style consistency suffers when work is split across multiple people or done ad hoc — a real problem for anyone trying to hold a brand/production standard. + +### Existing alternatives + +*What do people with this problem do today? (Remotion/DaVinci/Premiere themselves, a video editor they hire, just not editing multi-cam at all, Descript, etc.) Naming the real alternative — including "doing nothing" or "hiring a freelancer" — matters more than naming a direct competitor.* + +- DIY with free/prosumer tools: CapCut, Clipchamp, Riverside — free or cheap but buggy, slow, no automation for multi-cam jump cuts or transcript-driven editing. +- Paid editing services/studios: e.g. Poddster, ~$200+ for a full episode edit and another ~$200+ for 3 short-form clips — professional quality but expensive per-episode, doesn't scale with volume. +- In-house or freelance editor — costs money, and (per our own experience) still needs consistency/QA oversight. +- Doing nothing / publishing raw or minimally-edited footage — some early podcasters skip proper editing entirely due to cost/time, at the expense of quality. + +## 2. Customer Segments + +*Who exactly. Refine the "who else has this problem" guess from the hypothesis doc into something specific enough to recruit interviews against.* + +- **Segment 1 — primary, for canvas/testing purposes: B2B2B subcontracted editing.** Podcast studio rental companies (and similar production agencies) who want to offer editing services to their own clients without building/staffing that capability themselves. We do the editing; they route the work and own the client relationship. Currently in partnership negotiation with one such studio. This is the segment with a real competitive purchase decision (vs. Poddster, freelancers, in-house), so it's the one to write the UVP/pricing against. +- **Segment 2 — live, undecided priority: companies running podcasts for marketing/lead-gen.** May want done-for-you editing without touching the tool themselves. Likely converges with Segment 1 if they arrive via a studio partner; could also be a direct relationship. +- **Segment 3 — hypothesis, not yet engaged: solo/small podcast teams editing for themselves** (from the hypothesis doc's original "people with full-time jobs, editing on their own with tools like Riverside"). Relevant mainly if/when we pursue a self-serve product path later — not the near-term focus. +- **ragTech itself is not treated as a canvas customer segment** — we're the maker, not a market test, since we can't lose our own business to a competitor. But note: the actual day-to-day *operator* of the tool is a ragTech team member either way (our own episodes or Segment 1 client work) — the product has to work for that operator persona regardless of which segment's footage it is. + +### Early adopters + +*Within that segment, who would try something rough/unfinished first? These are who you interview and design-partner with first — not the mainstream of the segment.* + +- The specific podcast studio rental company currently in partnership talks — real, already in motion, highest priority to convert into a design partner for both the service and the underlying product. +- Other studio-rental or podcast-production agencies in the same category (similar to thelfgpod.com) who may have the same "we want to offer editing without building it ourselves" need. + +## 3. Unique Value Proposition + +*One sentence. What do you offer that's different/better, and why should a target customer care in the first 10 seconds?* + +- Fast, consistent multi-cam podcast editing — delivered in hours, not days — powered by efficient local AI instead of expensive cloud-token models. +- *(For Segment 1 specifically: reliable, affordable subcontracted editing capacity a studio can route client work to without building an editing team themselves.)* + +## 4. Solution + +*Top 3 features that address the top problems above. Resist listing everything in the RFC's build order — this is the customer-facing subset, not the technical roadmap.* + +1. Automated multi-cam jump-cut editing driven by the transcript (sync, camera switching, cuts) — removes the most repetitive manual work. +2. Fast native rendering — turns a many-hour render into under an hour, enabling same-day turnaround for subcontracted client work. +3. Built-in short-form/clip repurposing from the same long-form source — covers the "another full day" of clipping work in one pipeline, not a separate tool/pass. + +## 5. Channels + +*How would target customers actually find/reach this? (Podcast communities, YouTube creator forums, word of mouth from ragTech's own audience, etc.)* + +- B2B2B: direct relationship-building with podcast studio rental companies / production agencies — the current partnership is the first channel, not a hypothetical one. +- ragTech's own audience/network — credibility as working podcasters who built this to solve our own problem. +- Podcast creator communities — relevant mainly if/when Segment 2 or 3 becomes the near-term focus. + +## 6. Revenue Streams + +*If this becomes a product — subscription, one-time license, usage-based, or genuinely undecided. It's fine to write "undecided" but write down the options you're actually considering.* + +- Near-term (primary): per-episode / per-project subcontracted editing fee, paid by the studio partner (or their client, routed through the studio). Pricing itself is undecided — Poddster's ~$200/episode + ~$200/3 shorts is our current market reference point, not our price. +- Undecided within the partnership negotiation itself: flat fee vs. revenue share with the studio partner. +- Possible future, not v1: subscription or usage-based pricing if/when a self-serve product path opens up. + +## 7. Cost Structure + +*Rough — engineering time (the big one right now), infra/hosting if any, model/API costs if you keep any cloud-dependent ML features.* + +- Engineering time — dominant cost right now (the native rewrite itself). +- Editor labor time — once doing paid client work, our own team's time editing client episodes becomes a real cost/opportunity-cost, not just "our own free time" as it is today. +- Compute — local/on-device AI models (whisper.cpp, ONNX) avoid ongoing cloud API/token costs; this is a deliberate cost-structure choice, not just a UVP talking point. +- No hosting/infra cost currently — this is a native desktop tool operated internally, not a hosted service. + +## 8. Key Metrics + +*The few numbers that would tell you this is working. (E.g. time-to-render-an-episode, weekly active editors, retention after first project, NPS.)* + +- Render time per episode (RFC target: under 1 hour, down from up to 8). +- Edit turnaround time per client project — raw footage to delivered edit. +- Number of client teams/projects handled per week (12-month target: 10). +- Rework/revision rate per client project — a consistency proxy. +- Cost per episode edited (labor + compute) vs. revenue per episode — unit economics of the service business. + +## 9. Unfair Advantage + +*Something competitors can't easily copy. Be honest — for a v1 this might legitimately be "nothing yet" or "our own dogfooding/domain expertise as working podcasters," which is a real but modest advantage.* + +- We are working podcasters *and* full-time software engineers — deep, lived domain expertise in the exact workflow we're building for, plus the ability to build the fix ourselves rather than commission it or wait for someone else to. That combination is rare: most people who feel this pain can't build the solution, and most people who could build it don't feel the pain. +- Local/efficient AI architecture (vs. expensive cloud-token competitors) is a real, structural cost advantage tied directly to the native rewrite's technical decisions — not just brand positioning. +- An inbound B2B2B relationship already in motion — real distribution advantage vs. competitors who'd have to cold-sell into studio partnerships from scratch. + +--- + +## Revision log + +*Every time you revise a box based on real evidence (an interview, a churned assumption), note it here with the date and what changed. This is what makes the canvas a living hypothesis document instead of a one-time exercise.* + +| Date | Box changed | What changed | Why (evidence) | +|------|-------------|--------------|-----------------| +| | | | | diff --git a/product/02-stakeholder-map.md b/product/02-stakeholder-map.md new file mode 100644 index 0000000..ae70041 --- /dev/null +++ b/product/02-stakeholder-map.md @@ -0,0 +1,35 @@ +# Stakeholder Map + +Quick exercise — mostly internal right now. Purpose: know who has decision rights over scope before you start external conversations that create expectations you can't unilaterally act on. + +## Core team + +*Confirm roles and decision rights explicitly — don't assume everyone agrees who owns what.* + +| Name | Role | Decision rights (what can they say yes/no to alone?) | Primary concern | +|------|------|--------------------------------------------------------|------------------| +| Natasha | Software Engineer | | | +| Saloni | Software Developer | | | +| Victoria | Solutions Engineer | | | + +## Extended / occasional stakeholders + +*Advisors, investors, contractors, anyone with influence but not day-to-day involvement. If none yet, write "none yet" — don't leave it blank, since that's a different fact than "not applicable."* + +| Name/Group | Relationship | Influence level | What they need to be kept informed of | +|------------|---------------|------------------|------------------------------------------| +| | | | | + +## Prospective external stakeholders (product path) + +*People not yet involved but relevant if this becomes a product — early design partners you already know, a target customer community, potential early beta users.* + +| Name/Group | Relationship | Why they matter | Next step to engage them | +|------------|---------------|------------------|----------------------------| +| | | | | + +## Interest/influence grid (optional, quick) + +*Plot each stakeholder above: high-influence/high-interest people need active management; high-influence/low-interest need to be kept satisfied without over-communicating; low-influence/high-interest are good candidates for early interviews/beta.* + +- diff --git a/product/03-customer-interviews/README.md b/product/03-customer-interviews/README.md new file mode 100644 index 0000000..f321279 --- /dev/null +++ b/product/03-customer-interviews/README.md @@ -0,0 +1,39 @@ +# Customer Interviews + +Goal: test the riskiest assumptions in `01-lean-canvas.md` against people **outside ragTech**. You already know your own pain intimately — the open question is whether it generalizes, to whom, and whether it's painful enough that someone would switch tools or pay. + +## Who to recruit + +Pull from the Lean Canvas's "Early adopters" box. Aim for 5-8 people across your target segment(s) for a first round — enough to spot patterns, not so many you're stalling on research instead of building. Prioritize people who currently do multi-camera/multi-speaker video editing (other podcasters, video-first creators, small content teams) over people who don't edit video at all. + +## How to run it (Mom Test style) + +The #1 failure mode is asking questions that let people be polite instead of honest. Ground rules: + +- Ask about **specific past behavior**, not hypotheticals. "Walk me through the last time you edited an episode" beats "would you use a tool that does X?" +- Don't pitch the product during the interview. You're gathering evidence, not selling. +- Follow the pain, not your solution. If they don't bring up something close to your hypothesized problem unprompted, that's a real signal — don't lead them to it. +- Ask about money/switching concretely: "What do you currently pay for editing (tool, freelancer, your own time)?" / "What would have to be true for you to switch from what you use today?" + +## Starter question list + +*Adapt per segment, but keep the shape — concrete and past-tense.* + +1. Walk me through the last time you produced an episode/video, start to finish. Where did the time actually go? +2. What was the most annoying or frustrating part of that process? +3. What tools do you currently use for editing/production? What do you like and dislike about each? +4. Have you tried to fix or work around [the specific pain you're probing]? What did you try? Did it work? +5. What do you currently pay for tools, freelance editing, or your own time on this — roughly? +6. Who else on your team is involved in this workflow, and where does it get handed off? +7. If you could wave a magic wand and fix one thing about this process, what would it be? *(Ask this last — it's the closest thing to a leading question on this list.)* + +## Logistics + +- Record with permission; store raw recordings/transcripts outside git if they contain identifying info (see `.gitignore` note below). +- Transcribe (even roughly) within a day or two — memory of nuance fades fast. +- After each interview, spend 10 minutes writing 3-5 bullet takeaways in `synthesis.md` before moving to the next one. Don't batch this to the end. + +## Files in this folder + +- `transcripts/` — raw or lightly-cleaned transcripts. **Redact/anonymize before committing** if names or company details are sensitive; consider keeping raw audio out of git entirely (see the repo's `.gitignore` conventions for runtime/generated artifacts). +- `synthesis.md` — the living summary: recurring themes, direct quotes worth keeping, surprises, and anything that contradicts the Lean Canvas. diff --git a/product/03-customer-interviews/synthesis.md b/product/03-customer-interviews/synthesis.md new file mode 100644 index 0000000..e1c7e02 --- /dev/null +++ b/product/03-customer-interviews/synthesis.md @@ -0,0 +1,37 @@ +# Interview Synthesis + +Update this after every interview — don't wait until the end of the round. The goal is a living document, not a final report. + +## Interview log + +| # | Date | Who (role/segment) | Link to transcript | +|---|------|----------------------|----------------------| +| 1 | | | | + +## Recurring themes + +*Add a theme as soon as you hear it in 2+ interviews. One bullet per theme, with which interview #s support it.* + +- + +## Direct quotes worth keeping + +*Verbatim, with interview # and enough context to know what prompted it. These become evidence in `06-opportunities/` and later in engineering issues — a real quote is worth more than a paraphrase.* + +- + +## Surprises / things that contradicted our assumptions + +*Be honest about these — they're the most valuable findings, not the ones that confirm what you already believed.* + +- + +## Signals on willingness to pay / switch + +*What did people actually say about current spend, and what would make them switch? Not what you inferred — what they said.* + +- + +## Open questions for the next round + +- diff --git a/product/03-customer-interviews/transcripts/README.md b/product/03-customer-interviews/transcripts/README.md new file mode 100644 index 0000000..b97d340 --- /dev/null +++ b/product/03-customer-interviews/transcripts/README.md @@ -0,0 +1,22 @@ +# Transcripts + +Raw or lightly-cleaned interview transcripts go here, one file per interview (e.g. `01-jane-doe-podcaster.md`). + +**Before committing:** redact names/companies if there's any sensitivity, or keep this folder's contents out of git entirely and store transcripts in a shared drive instead — the repo's `.gitignore` already excludes everything in this folder except this README so that's the default. Remove the ignore rule for this path if you decide raw transcripts are fine to commit as-is. + +Suggested per-file format: + +```markdown +# Interview 01 — [role/segment, not necessarily full name if sensitive] + +**Date:** +**Segment:** (which Lean Canvas customer segment) + +## Notes/Transcript + +... + +## Immediate takeaways (write this within an hour of the call) + +- +``` diff --git a/product/04-journey-map.md b/product/04-journey-map.md new file mode 100644 index 0000000..bf1cd46 --- /dev/null +++ b/product/04-journey-map.md @@ -0,0 +1,38 @@ +# Customer Journey Map + +Build this **after** the first round of interviews (`03-customer-interviews/synthesis.md`) — it should reflect what a target customer's workflow actually looks like, not just ragTech's own pipeline. Use your own dogfooding (the technical stages are already documented in `docs/rfcs/0001-native-desktop-rewrite.md` and the root `CLAUDE.md`) as one data point, not the only one. + +Map one representative persona/segment at a time — if you have multiple segments with meaningfully different workflows (e.g. solo creator vs. small team), do a separate map for each rather than averaging them into a mushy composite. + +## Persona / segment for this map + +*Which persona from `05-personas.md` (or interview segment) this map represents.* + +- + +## Stages + +*Adapt these to what interviews actually revealed — this is a reasonable starting skeleton based on ragTech's own pipeline, but a target customer's real workflow may have different or extra stages (e.g. scripting/planning before recording, client review/approval after editing).* + +| Stage | What they're doing | What they're thinking/feeling | Pain points (with evidence) | Time cost | Opportunity | +|-------|----------------------|----------------------------------|-------------------------------|-----------|--------------| +| Record | | | | | | +| Sync (multi-angle/audio) | | | | | | +| Transcribe | | | | | | +| Diarize/label speakers | | | | | | +| Edit transcript (cuts, corrections) | | | | | | +| Camera/framing decisions | | | | | | +| Render/export | | | | | | +| Publish/distribute | | | | | | + +## Biggest pain point overall + +*The single stage/moment that came up most often or most intensely across interviews. This is usually where the first opportunity doc should focus.* + +- + +## Evidence trail + +*For each pain point above, link back to the specific interview # and quote in `03-customer-interviews/synthesis.md` that supports it. If a row has no evidence trail, treat it as your own assumption, not a validated finding — mark it as such.* + +- diff --git a/product/05-personas.md b/product/05-personas.md new file mode 100644 index 0000000..c6d7e79 --- /dev/null +++ b/product/05-personas.md @@ -0,0 +1,33 @@ +# Personas + +Keep this light — 1-2 primary personas, distilled from actual interview clusters in `03-customer-interviews/synthesis.md`, not invented from imagination. If you don't have enough interviews yet to cluster, don't force a persona — write "not enough evidence yet" rather than a made-up one; a fabricated persona is worse than none because it feels authoritative without being grounded. + +## Persona template (copy per persona) + +### [Name] — [one-line descriptor, e.g. "solo podcast host editing their own multi-cam episodes"] + +**Segment:** (from Lean Canvas customer segments) + +**Role/context:** + +**Goals:** *What are they actually trying to accomplish — not "use good software" but the real outcome (e.g. "publish weekly without burning a full day on editing").* + +**Frustrations:** *Pull directly from interview quotes where possible.* + +**Representative quote:** *A real quote from `03-customer-interviews/`, not a paraphrase.* + +**Current tools:** + +**Tech comfort level:** + +**What would make them switch to a new tool:** + +--- + +## Persona 1 + +*(fill in using the template above)* + +## Persona 2 + +*(only add if interviews genuinely surfaced a second distinct cluster — don't force it)* diff --git a/product/06-opportunities/README.md b/product/06-opportunities/README.md new file mode 100644 index 0000000..d59f5c2 --- /dev/null +++ b/product/06-opportunities/README.md @@ -0,0 +1,31 @@ +# Opportunities + +One file per opportunity, named `NN-short-slug.md` (e.g. `01-faster-multicam-render.md`). This is the direct bridge from product discovery to engineering — every opportunity doc here should become one or more epics/issues in the new repo, and every engineering issue in the new repo should be traceable back to one of these. + +**Rule: if you can't fill in the Evidence section, it's not ready to become an opportunity doc yet.** Go back to the journey map or interviews first. An opportunity based only on "we personally find this annoying" is a hypothesis, not a validated opportunity — that's fine, just label it as such rather than skipping the step. + +## Template (copy into a new file per opportunity) + +```markdown +# Opportunity: [short name] + +## Problem statement +One or two sentences. What's broken, for whom, framed from the customer's side (not "our renderer is slow" but "editors wait N hours before they can review a cut"). + +## Evidence +- Journey map stage(s): (link to `../04-journey-map.md` row) +- Interview quotes/findings: (link to specific interview #s in `../03-customer-interviews/synthesis.md`) +- If evidence is currently only "our own dogfooding, not yet validated externally" — say so explicitly. + +## Proposed direction +What we think the solution looks like, at a product level (not implementation detail — that's what the engineering issue is for). + +## Success metric +How we'd know this opportunity was actually addressed. Tie to a Lean Canvas Key Metric if possible. + +## Related RFC / technical context +Link to the relevant section of `docs/rfcs/0001-native-desktop-rewrite.md` if this opportunity has technical implications already scoped there. + +## Status +Hypothesis / Validated / In progress / Shipped +``` diff --git a/product/README.md b/product/README.md new file mode 100644 index 0000000..9bb7c6a --- /dev/null +++ b/product/README.md @@ -0,0 +1,30 @@ +# Product Founding Docs + +This directory is the product/discovery half of the native-rewrite planning work — the companion to [`docs/rfcs/0001-native-desktop-rewrite.md`](../docs/rfcs/0001-native-desktop-rewrite.md), which is the *technical* half. That RFC justifies the rewrite from engineering pain (Remotion's render ceiling, hardware inconsistency). This directory exists to justify it from **customer** pain, and to build the artifacts that turn "we personally find this painful" into engineering issues grounded in evidence beyond the three of us. + +It's designed to be portable: when the new repo is created, copy this whole `product/` folder over as-is. Nothing in here depends on the current codebase. + +## Why this exists + +ragTech's editing pipeline started as an internal tool. This rewrite is being scoped as if it might become a product other podcasters/video teams use — which means before writing engineering issues, we should know: is the pain we feel actually general, who else has it, would they pay, and what does their current workflow actually look like (not just ours). + +## Recommended order + +Do these roughly in sequence — each one narrows/grounds the next. Don't skip straight to interviews or a PRD without the earlier steps; you'll end up validating assumptions you never stated. + +| # | File | Purpose | Time to first draft | +|---|------|---------|---------------------| +| 1 | [`00-problem-hypothesis.md`](00-problem-hypothesis.md) | Your own claim, before any research — what's broken, who else might have it, why now | 30–60 min | +| 2 | [`01-lean-canvas.md`](01-lean-canvas.md) | Turn the hypothesis into 9 falsifiable claims on one page | 45 min | +| 3 | [`02-stakeholder-map.md`](02-stakeholder-map.md) | Who has a say, who's affected, who needs to be kept informed | 20 min | +| 4 | [`03-customer-interviews/`](03-customer-interviews/) | Test the riskiest Lean Canvas boxes against real people outside ragTech | 1–2 weeks (5–8 interviews) | +| 5 | [`04-journey-map.md`](04-journey-map.md) | Map a target customer's current workflow, pain-annotated, using interview findings + our own dogfooding | after interviews | +| 6 | [`05-personas.md`](05-personas.md) | 1–2 lightweight personas distilled from interview clusters | after interviews | +| 7 | [`06-opportunities/`](06-opportunities/) | One doc per opportunity surfaced above — the direct bridge to epics/issues in the new repo | ongoing | + +## What "done" looks like before writing the first engineering issue in the new repo + +- Lean Canvas has been revised at least once based on real interview feedback (not just your own assumptions). +- At least 5 interviews with people outside ragTech, synthesized (not just raw transcripts sitting unread). +- A journey map that names specific pain points with evidence (a quote, a timing observation), not just "editing is slow." +- Each opportunity doc traces back to a specific journey-map pain point or interview finding — if you can't point to the evidence, it's not ready to become an epic yet. From 3ee86a8060f5b5de735ff390e3eac8a62753f3fb Mon Sep 17 00:00:00 2001 From: Natasha Ann Lum Date: Fri, 10 Jul 2026 14:59:16 +0800 Subject: [PATCH 2/7] docs: reconcile Lean Canvas with poddedit pricing/negotiation collateral Names the product (poddedit, distinct from ragTech the team/podcast) and replaces "undecided" placeholders in Existing Alternatives, Customer Segments, Revenue Streams, Cost Structure, and Key Metrics with the concrete tiers, three named segments, and financial targets already drafted on the editing-services branch for the podcast-studio negotiation. Logged in the canvas's revision log as evidence-based, not yet customer-validated. Co-Authored-By: Claude Sonnet 5 --- product/01-lean-canvas.md | 25 ++++++++++++++----------- product/README.md | 8 ++++++-- 2 files changed, 20 insertions(+), 13 deletions(-) diff --git a/product/01-lean-canvas.md b/product/01-lean-canvas.md index a0c7c29..37b7af5 100644 --- a/product/01-lean-canvas.md +++ b/product/01-lean-canvas.md @@ -17,7 +17,7 @@ Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 b *What do people with this problem do today? (Remotion/DaVinci/Premiere themselves, a video editor they hire, just not editing multi-cam at all, Descript, etc.) Naming the real alternative — including "doing nothing" or "hiring a freelancer" — matters more than naming a direct competitor.* - DIY with free/prosumer tools: CapCut, Clipchamp, Riverside — free or cheap but buggy, slow, no automation for multi-cam jump cuts or transcript-driven editing. -- Paid editing services/studios: e.g. Poddster, ~$200+ for a full episode edit and another ~$200+ for 3 short-form clips — professional quality but expensive per-episode, doesn't scale with volume. +- Paid editing services/studios: e.g. Poddster (Singapore) — S$273 recording + S$280 standard episode edit + S$453 manual-subtitle add-on + S$280 for 3 highlight clips = S$1,286/cycle for a comparable bundled outcome; line-item add-on pricing, not bundled. Full comparison in `docs/PRICING_TIERS.md` (`editing-services` branch). - In-house or freelance editor — costs money, and (per our own experience) still needs consistency/QA oversight. - Doing nothing / publishing raw or minimally-edited footage — some early podcasters skip proper editing entirely due to cost/time, at the expense of quality. @@ -25,10 +25,13 @@ Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 b *Who exactly. Refine the "who else has this problem" guess from the hypothesis doc into something specific enough to recruit interviews against.* -- **Segment 1 — primary, for canvas/testing purposes: B2B2B subcontracted editing.** Podcast studio rental companies (and similar production agencies) who want to offer editing services to their own clients without building/staffing that capability themselves. We do the editing; they route the work and own the client relationship. Currently in partnership negotiation with one such studio. This is the segment with a real competitive purchase decision (vs. Poddster, freelancers, in-house), so it's the one to write the UVP/pricing against. -- **Segment 2 — live, undecided priority: companies running podcasts for marketing/lead-gen.** May want done-for-you editing without touching the tool themselves. Likely converges with Segment 1 if they arrive via a studio partner; could also be a direct relationship. -- **Segment 3 — hypothesis, not yet engaged: solo/small podcast teams editing for themselves** (from the hypothesis doc's original "people with full-time jobs, editing on their own with tools like Riverside"). Relevant mainly if/when we pursue a self-serve product path later — not the near-term focus. -- **ragTech itself is not treated as a canvas customer segment** — we're the maker, not a market test, since we can't lose our own business to a competitor. But note: the actual day-to-day *operator* of the tool is a ragTech team member either way (our own episodes or Segment 1 client work) — the product has to work for that operator persona regardless of which segment's footage it is. +Three named segments (aligned with `docs/EDITING_PACKAGES.md` on the `editing-services` branch): + +- **Segment 1 — primary near-term, real deal in motion: Podcast Studios and Agency Partners (B2B2B).** Studios/agencies running production for multiple podcast clients who want to expand editing capacity/margin without building an in-house team. We fulfill; they route work and own the client relationship. Currently negotiating with one such studio — pricing tiers (Essential/Growth/Premium, S$179-S$749) are being proposed as what the studio can offer *its* clients, with poddedit fulfilling behind the scenes. This is the segment with a real, live competitive purchase decision, so it's the one the UVP is written against. +- **Segment 2 — live: Independent and Creator-Led Podcasts.** Solo hosts/small teams publishing consistently on YouTube/Spotify/TikTok, need multi-format repurposing without a large time investment. Could be reached directly (tiered self-serve packages) or via a studio/agency. +- **Segment 3 — live: Business and Thought-Leadership Podcasts.** Founder-led/B2B/company podcasts using episodes for brand and demand-gen; want polished output without managing in-house post-production. Likely willing to pay more for reliability/speed than Segment 2, per willingness-to-pay analysis in `docs/FINANCIAL_PROJECTIONS.md`. +- Segment priority isn't fully locked — Segment 1 is the immediate focus because a deal is actively in motion, but 2 and 3 are real, named, and not just hypotheses waiting on Segment 1 to fall through. +- **ragTech itself is not treated as a canvas customer segment** — we're the maker, not a market test, since we can't lose our own business to a competitor. But note: the actual day-to-day *operator* of the tool is a ragTech team member either way (our own episodes or client work) — poddedit has to work for that operator persona regardless of which segment's footage it is. ### Early adopters @@ -64,16 +67,16 @@ Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 b *If this becomes a product — subscription, one-time license, usage-based, or genuinely undecided. It's fine to write "undecided" but write down the options you're actually considering.* -- Near-term (primary): per-episode / per-project subcontracted editing fee, paid by the studio partner (or their client, routed through the studio). Pricing itself is undecided — Poddster's ~$200/episode + ~$200/3 shorts is our current market reference point, not our price. -- Undecided within the partnership negotiation itself: flat fee vs. revenue share with the studio partner. -- Possible future, not v1: subscription or usage-based pricing if/when a self-serve product path opens up. +- **Primary: per-episode tiered packages.** Essential (S$179-229) / Growth (S$279-399) / Premium (S$499-749), scaled by episode length, plus add-ons (extra clip +S$30, extra carousel +S$50, rush +25%, chapter markers S$29, custom branding S$299++, annual retainer = 1 month free). Full detail in `docs/PRICING_TIERS.md` and `docs/EDITING_PACKAGES.md` (`editing-services` branch). +- Deal-specific and still to be settled in the studio negotiation itself: whether the studio pays these tiers directly as a wholesale/fulfillment rate (flat fee or rev-share with the studio's own markup on top) vs. some other split — this doc's `docs/EDITING_PACKAGES.md` presents the tiers as what a studio could offer *its own clients*, not necessarily poddedit's price to the studio itself. +- Possible future, not v1: subscription or usage-based self-serve pricing if a direct-to-creator/product path opens up (Segments 2/3 without a studio intermediary). ## 7. Cost Structure *Rough — engineering time (the big one right now), infra/hosting if any, model/API costs if you keep any cloud-dependent ML features.* - Engineering time — dominant cost right now (the native rewrite itself). -- Editor labor time — once doing paid client work, our own team's time editing client episodes becomes a real cost/opportunity-cost, not just "our own free time" as it is today. +- Editor labor time — the real constraint once doing paid client work. `docs/FINANCIAL_PROJECTIONS.md` estimates hands-on time per tier (Essential ~1.0h, Growth ~2.0h, Premium ~4.0h) implying effective hourly returns of S$156-204/h across tiers even before the render-speed improvements this rewrite targets — meaning faster rendering directly increases effective hourly rate, not just "nice to have." - Compute — local/on-device AI models (whisper.cpp, ONNX) avoid ongoing cloud API/token costs; this is a deliberate cost-structure choice, not just a UVP talking point. - No hosting/infra cost currently — this is a native desktop tool operated internally, not a hosted service. @@ -83,7 +86,7 @@ Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 b - Render time per episode (RFC target: under 1 hour, down from up to 8). - Edit turnaround time per client project — raw footage to delivered edit. -- Number of client teams/projects handled per week (12-month target: 10). +- Episode volume per person/month — `docs/FINANCIAL_PROJECTIONS.md` Scenario B (Growth-led) targets ~17 episodes/month/person at S$5,658/month; team-of-3 target S$16,500-18,000/month (~51 episodes/month combined). These are the concrete version of the 12-month "10 client teams weekly" success metric in `00-problem-hypothesis.md` — worth reconciling the two into one number once the studio deal's actual terms are known. - Rework/revision rate per client project — a consistency proxy. - Cost per episode edited (labor + compute) vs. revenue per episode — unit economics of the service business. @@ -103,4 +106,4 @@ Reference: [Ash Maurya's Lean Canvas](https://leanstack.com/lean-canvas) — 9 b | Date | Box changed | What changed | Why (evidence) | |------|-------------|--------------|-----------------| -| | | | | +| 2026-07-10 | Existing Alternatives, Customer Segments, Revenue Streams, Cost Structure, Key Metrics | Replaced "undecided" placeholders with concrete pricing tiers, three named segments, and financial targets | Reconciled with pre-existing negotiation collateral on the `editing-services` branch (`docs/PRICING_TIERS.md`, `EDITING_PACKAGES.md`, `FINANCIAL_PROJECTIONS.md`), written by Natasha to negotiate the podcast-studio partnership offer. Not yet validated by customer interviews — still a hypothesis, just a more concrete one. | diff --git a/product/README.md b/product/README.md index 9bb7c6a..97c5bcf 100644 --- a/product/README.md +++ b/product/README.md @@ -1,6 +1,10 @@ -# Product Founding Docs +# Product Founding Docs — poddedit -This directory is the product/discovery half of the native-rewrite planning work — the companion to [`docs/rfcs/0001-native-desktop-rewrite.md`](../docs/rfcs/0001-native-desktop-rewrite.md), which is the *technical* half. That RFC justifies the rewrite from engineering pain (Remotion's render ceiling, hardware inconsistency). This directory exists to justify it from **customer** pain, and to build the artifacts that turn "we personally find this painful" into engineering issues grounded in evidence beyond the three of us. +**Naming note:** ragTech is the podcast/team. **poddedit** is the name of the product being built — the native rewrite described in the technical RFC, plus the editing-service business built on top of it. Use "poddedit" when referring to the product itself in these docs; "ragTech" refers to the team/podcast specifically. + +This directory is the product/discovery half of the poddedit planning work — the companion to [`docs/rfcs/0001-native-desktop-rewrite.md`](../docs/rfcs/0001-native-desktop-rewrite.md), which is the *technical* half. That RFC justifies the rewrite from engineering pain (Remotion's render ceiling, hardware inconsistency). This directory exists to justify it from **customer** pain, and to build the artifacts that turn "we personally find this painful" into engineering issues grounded in evidence beyond the three of us. + +Related, deal-specific collateral currently lives on the `editing-services` branch (`docs/PRICING_TIERS.md`, `docs/EDITING_PACKAGES.md`, `docs/FINANCIAL_PROJECTIONS.md`) — written to negotiate the podcast-studio partnership. The concrete pricing/market figures from that work have been folded into `01-lean-canvas.md` below; see that branch for the full negotiation-specific detail. It's designed to be portable: when the new repo is created, copy this whole `product/` folder over as-is. Nothing in here depends on the current codebase. From 0db59de5d72ed3de59d516a7e8031cbeefbdd7d5 Mon Sep 17 00:00:00 2001 From: Natasha Ann Date: Sun, 12 Jul 2026 16:57:03 +0800 Subject: [PATCH 3/7] chore(product): moved product docs to /docs --- {product => docs/product}/00-problem-hypothesis.md | 0 {product => docs/product}/01-lean-canvas.md | 0 {product => docs/product}/02-stakeholder-map.md | 0 {product => docs/product}/03-customer-interviews/README.md | 0 {product => docs/product}/03-customer-interviews/synthesis.md | 0 .../product}/03-customer-interviews/transcripts/README.md | 0 {product => docs/product}/04-journey-map.md | 0 {product => docs/product}/05-personas.md | 0 {product => docs/product}/06-opportunities/README.md | 0 {product => docs/product}/README.md | 0 10 files changed, 0 insertions(+), 0 deletions(-) rename {product => docs/product}/00-problem-hypothesis.md (100%) rename {product => docs/product}/01-lean-canvas.md (100%) rename {product => docs/product}/02-stakeholder-map.md (100%) rename {product => docs/product}/03-customer-interviews/README.md (100%) rename {product => docs/product}/03-customer-interviews/synthesis.md (100%) rename {product => docs/product}/03-customer-interviews/transcripts/README.md (100%) rename {product => docs/product}/04-journey-map.md (100%) rename {product => docs/product}/05-personas.md (100%) rename {product => docs/product}/06-opportunities/README.md (100%) rename {product => docs/product}/README.md (100%) diff --git a/product/00-problem-hypothesis.md b/docs/product/00-problem-hypothesis.md similarity index 100% rename from product/00-problem-hypothesis.md rename to docs/product/00-problem-hypothesis.md diff --git a/product/01-lean-canvas.md b/docs/product/01-lean-canvas.md similarity index 100% rename from product/01-lean-canvas.md rename to docs/product/01-lean-canvas.md diff --git a/product/02-stakeholder-map.md b/docs/product/02-stakeholder-map.md similarity index 100% rename from product/02-stakeholder-map.md rename to docs/product/02-stakeholder-map.md diff --git a/product/03-customer-interviews/README.md b/docs/product/03-customer-interviews/README.md similarity index 100% rename from product/03-customer-interviews/README.md rename to docs/product/03-customer-interviews/README.md diff --git a/product/03-customer-interviews/synthesis.md b/docs/product/03-customer-interviews/synthesis.md similarity index 100% rename from product/03-customer-interviews/synthesis.md rename to docs/product/03-customer-interviews/synthesis.md diff --git a/product/03-customer-interviews/transcripts/README.md b/docs/product/03-customer-interviews/transcripts/README.md similarity index 100% rename from product/03-customer-interviews/transcripts/README.md rename to docs/product/03-customer-interviews/transcripts/README.md diff --git a/product/04-journey-map.md b/docs/product/04-journey-map.md similarity index 100% rename from product/04-journey-map.md rename to docs/product/04-journey-map.md diff --git a/product/05-personas.md b/docs/product/05-personas.md similarity index 100% rename from product/05-personas.md rename to docs/product/05-personas.md diff --git a/product/06-opportunities/README.md b/docs/product/06-opportunities/README.md similarity index 100% rename from product/06-opportunities/README.md rename to docs/product/06-opportunities/README.md diff --git a/product/README.md b/docs/product/README.md similarity index 100% rename from product/README.md rename to docs/product/README.md From 3469fd9aab3c26f8cdd51fc079ed75dc61b60299 Mon Sep 17 00:00:00 2001 From: Natasha Ann Date: Sun, 12 Jul 2026 21:40:22 +0800 Subject: [PATCH 4/7] docs(product): add opportunities --- docs/product/04-journey-map.md | 4 +- docs/product/05-personas.md | 30 ++++++++++-- .../06-opportunities/01-render-speed.md | 36 ++++++++++++++ .../02-hardware-inconsistency.md | 34 ++++++++++++++ .../03-edit-time-end-to-end.md | 47 +++++++++++++++++++ .../04-transcript-timing-accuracy.md | 35 ++++++++++++++ .../05-visual-transcript-editor.md | 42 +++++++++++++++++ .../06-multi-brand-support.md | 44 +++++++++++++++++ docs/product/README.md | 24 ++++++++-- 9 files changed, 286 insertions(+), 10 deletions(-) create mode 100644 docs/product/06-opportunities/01-render-speed.md create mode 100644 docs/product/06-opportunities/02-hardware-inconsistency.md create mode 100644 docs/product/06-opportunities/03-edit-time-end-to-end.md create mode 100644 docs/product/06-opportunities/04-transcript-timing-accuracy.md create mode 100644 docs/product/06-opportunities/05-visual-transcript-editor.md create mode 100644 docs/product/06-opportunities/06-multi-brand-support.md diff --git a/docs/product/04-journey-map.md b/docs/product/04-journey-map.md index bf1cd46..a831046 100644 --- a/docs/product/04-journey-map.md +++ b/docs/product/04-journey-map.md @@ -1,8 +1,8 @@ # Customer Journey Map -Build this **after** the first round of interviews (`03-customer-interviews/synthesis.md`) — it should reflect what a target customer's workflow actually looks like, not just ragTech's own pipeline. Use your own dogfooding (the technical stages are already documented in `docs/rfcs/0001-native-desktop-rewrite.md` and the root `CLAUDE.md`) as one data point, not the only one. +Build this from ragTech's own pipeline experience first — the technical stages are already documented in `docs/rfcs/0001-native-desktop-rewrite.md` and the root `CLAUDE.md`, and the time costs and failure modes are directly observed. Label each pain point's evidence source honestly (internal dogfooding, RFC analysis, financial modeling). When external interviews happen (Phase 2 per `README.md`), revise the relevant rows and update the evidence trail — don't wait to fill this in. -Map one representative persona/segment at a time — if you have multiple segments with meaningfully different workflows (e.g. solo creator vs. small team), do a separate map for each rather than averaging them into a mushy composite. +Map one representative persona/segment at a time — if you have multiple segments with meaningfully different workflows (e.g. ragTech operator vs. external solo creator), do a separate map for each rather than averaging them into a mushy composite. ## Persona / segment for this map diff --git a/docs/product/05-personas.md b/docs/product/05-personas.md index c6d7e79..3722a51 100644 --- a/docs/product/05-personas.md +++ b/docs/product/05-personas.md @@ -1,6 +1,8 @@ # Personas -Keep this light — 1-2 primary personas, distilled from actual interview clusters in `03-customer-interviews/synthesis.md`, not invented from imagination. If you don't have enough interviews yet to cluster, don't force a persona — write "not enough evidence yet" rather than a made-up one; a fabricated persona is worse than none because it feels authoritative without being grounded. +Keep this light — 1-2 primary personas. During Phase 1 (internal dogfooding — see `README.md`), personas can be derived from direct observation of how the tool is used day-to-day, not just from interviews. Label the evidence source clearly. When interviews happen (Phase 2), revise with real quote evidence from `03-customer-interviews/synthesis.md`. + +A persona invented from imagination with no evidence trail is worse than none — it feels authoritative without being grounded. But a persona derived from months of direct, daily use of the tool is real evidence; don't withhold it just because it came from dogfooding rather than a formal interview. ## Persona template (copy per persona) @@ -26,8 +28,30 @@ Keep this light — 1-2 primary personas, distilled from actual interview cluste ## Persona 1 -*(fill in using the template above)* +### The ragTech Operator — "engineer-editor rotating through the pipeline biweekly" + +**Segment:** Not a canvas customer segment — this is the internal operator persona. The actual day-to-day user of the tool is a ragTech team member regardless of whether the footage is our own episode or a client's. poddedit has to work for this persona before it can work for anyone else. + +**Evidence source:** Internal dogfooding — direct observation from building and operating the pipeline since it was first prototyped. Not yet validated externally. + +**Role/context:** Full-time software engineer who rotates podcast editing duty every two weeks. Not a professional editor. Knows the codebase well enough to debug when something breaks, but the goal is that they shouldn't have to. + +**Goals:** Get from raw footage to delivered edit (long-form + short-form clips) without spending more than 2 hours of hands-on time, without the render blocking them from doing other work, and without the output quality depending on which machine or which team member ran it. + +**Frustrations:** +- Render takes 6–18 hours — can't review the cut until the next day. +- Pipeline behaves differently on different machines (Mac M2 vs Mac M3 vs Windows/NVIDIA) — unclear whether a difference in output is a bug or hardware. +- Short-form clipping is a full separate manual pass, not a first-class output of the same run. +- When something breaks, the failure mode is often silent (wrong encoder silently chosen, stale frame decode not surfaced as an error). + +**Current tools:** The internal ragTech pipeline (deckcreate repo) — the very thing being rewritten. + +**Tech comfort level:** High — can read and modify the codebase, run scripts, debug FFmpeg flags. But the tool should not require this; the goal is that a team member who is less deep in the codebase can operate it without debugging. + +**What would make them switch (or: what does "done" look like for this persona):** Render under 1 hour, consistent output across all three hardware targets, short-form clips as a first-class output, and silent failures surfaced as actual errors. + +--- ## Persona 2 -*(only add if interviews genuinely surfaced a second distinct cluster — don't force it)* +*(Only add when interviews or real client usage surfaces a second distinct operator type — e.g. an external editor operating the tool for client work who is not a ragTech engineer. Don't force it before that evidence exists.)* diff --git a/docs/product/06-opportunities/01-render-speed.md b/docs/product/06-opportunities/01-render-speed.md new file mode 100644 index 0000000..70f8344 --- /dev/null +++ b/docs/product/06-opportunities/01-render-speed.md @@ -0,0 +1,36 @@ +# Opportunity: Render speed + +## Problem statement + +Editors wait up to 6–18 hours after completing a transcript edit before they can review and deliver a final cut. At that turnaround, same-day delivery of client work is impossible regardless of how fast the editing itself goes — the render is the wall. + +## Evidence + +- Journey map stage: Render/export — no entries yet (pre-interviews), but the time cost is directly observed from ragTech's own pipeline. +- **Internal dogfooding (not yet externally validated):** Remotion renders via headless Chromium at a 2–5 fps ceiling with no GPU path. A 60-minute episode at 60fps takes 6–18 hours. Multi-camera angles compound this: each inactive angle's `OffthreadVideo` must stay decoding at `opacity:1` simultaneously or its decoder stalls and produces stale frames on switch — decode cost scales linearly with angle count. Source: [`docs/rfcs/0001-native-desktop-rewrite.md` §Context #1](../rfcs/0001-native-desktop-rewrite.md). +- Lean Canvas §8 Key Metrics: render time per episode is listed as a primary metric; RFC target is under 1 hour from up to 8. +- Lean Canvas §7 Cost Structure: "faster rendering directly increases effective hourly rate, not just 'nice to have'" — `docs/FINANCIAL_PROJECTIONS.md` estimates effective rates of S$156–204/h per tier even before render-speed improvement; that rate compounds upward when render time stops being a floor constraint on episode volume. +- 12-month success metric from [`00-problem-hypothesis.md`](../00-problem-hypothesis.md): "ragTech renders episodes in under 1 hour instead of 8." + +## Proposed direction + +Replace the Remotion/headless-Chromium render path with a native compositor that routes encode/decode through platform GPU APIs (VideoToolbox on Mac, NVENC/NVDEC on Windows/Linux) and uses hardware-accelerated multi-angle decoding. Preview and final render share the same compositor code path — no second implementation that can drift. This removes the browser-engine ceiling entirely rather than tuning around it. + +## Success metric + +- Render time for a 60-minute, 3-angle episode at 60fps: under 1 hour on Mac M2, Mac M3, and Windows/NVIDIA — verified across all three hardware targets, not just the fastest machine. +- Lean Canvas §8 tie-in: "render time per episode" drops from the current 6–18h range into the sub-1h range across the full hardware matrix. +- Unit economics tie-in: effective hourly rate per episode tier (S$156–204/h baseline from `docs/FINANCIAL_PROJECTIONS.md`) should improve meaningfully once render time is no longer a floor constraint on daily episode volume. + +## Related RFC / technical context + +- [`docs/rfcs/0001-native-desktop-rewrite.md` §Context #1](../rfcs/0001-native-desktop-rewrite.md) — Remotion bottleneck analysis; multi-angle decode cost. +- RFC §Decision #1 — Rust chosen over C++ specifically for this render/compositor path. +- RFC §Decision #2 — Native egui/wgpu GUI (not Tauri) so preview shares the compositor rather than duplicating it. +- RFC §Decision #4 — Render/compositing engine is the first thing to port; hardware/encoder selection ported alongside it. +- RFC §Build Order Step 1 — Validate against real-episode fixtures with golden-frame diffing across Mac M2, Mac M3, and Windows/NVIDIA before trusting it. +- RFC §Open Questions — FFmpeg vs gstreamer-rs binding choice is unresolved; a spike must resolve this before the render engine epic begins. + +## Status + +Hypothesis — internal dogfooding only. Not yet validated against external customers. Priority: P0 — this is the blocker for the service business's unit economics and the primary motivation for the rewrite. diff --git a/docs/product/06-opportunities/02-hardware-inconsistency.md b/docs/product/06-opportunities/02-hardware-inconsistency.md new file mode 100644 index 0000000..2307c74 --- /dev/null +++ b/docs/product/06-opportunities/02-hardware-inconsistency.md @@ -0,0 +1,34 @@ +# Opportunity: Hardware-consistent output + +## Problem statement + +The pipeline produces different output depending on which machine runs it — a Windows machine with an NVIDIA GPU silently falls back to software encoding, while a Mac with VideoToolbox does not. Editors cannot guarantee that a render on one machine matches a render on another, which is a hidden quality risk when work rotates across team members or hardware. + +## Evidence + +- Journey map stage: Render/export — no entries yet (pre-interviews), but inconsistency is directly observed from ragTech's own cross-machine experience. +- **Internal dogfooding (not yet externally validated):** + - `scripts/config/hardware.ts` derives `encoderProfile` from `process.platform`/`process.arch` string-matching only: `supportsVideoToolbox = platform === 'darwin'`, `supportsCuda = platform === 'linux' && arch === 'x64'`. A Windows machine with an NVIDIA GPU never resolves to `nvenc` — it silently falls back to software `libx264`. The type's own doc comment confirms: "No FFmpeg integration — encoder flags are recorded but not applied here." Source: [RFC §Context #2](../rfcs/0001-native-desktop-rewrite.md). + - FFmpeg encoder selection is scattered across scripts: only `scripts/sync/AudioSyncer.js` branches on platform; `conform-to-raw.js`, `cut-preview.js`, `optimize-for-remotion.js`, and `transcode-proxy.js` hardcode `libx264` unconditionally. No `h264_nvenc`, `-hwaccel cuda`, or `scale_cuda`/`scale_metal` flags exist anywhere in the repo — NVENC/NVDEC is entirely unimplemented. Source: [RFC §Context #3](../rfcs/0001-native-desktop-rewrite.md). +- The current team has Mac M2, Mac M3, and a Windows/NVIDIA machine — all three are real, live hardware targets, and inconsistency between them is already observable. + +## Proposed direction + +Replace the platform-string heuristic with real hardware capability probing at startup (query the actual GPU/codec hardware, not just OS name), and consolidate the currently-scattered per-file encoder branching into a single `encoderProfile → ffmpeg args` mapping. Every script calls this one place; no script decides its own encoder. Implement NVENC/NVDEC paths so Windows/NVIDIA is a first-class target, not a silent fallback. + +## Success metric + +- A render of the same episode fixture on Mac M2, Mac M3, and Windows/NVIDIA produces output that passes a PSNR/hash diff gate — frame-level parity, not approximate visual similarity. +- `hardware.ts`-equivalent capability probe correctly identifies VideoToolbox on Mac and NVENC on Windows/NVIDIA without false negatives, verified by running on all three machines in CI. +- No script in the new repo hardcodes an encoder; all route through the single encoder-profile mapping. + +## Related RFC / technical context + +- [RFC §Context #2](../rfcs/0001-native-desktop-rewrite.md) — `hardware.ts` stub analysis. +- [RFC §Context #3](../rfcs/0001-native-desktop-rewrite.md) — scattered FFmpeg encoder selection; NVENC unimplemented. +- RFC §Decision #4 — Hardware/encoder selection is ported alongside the render engine (not a later phase); the two are tightly coupled. +- RFC §Verification §6 — Cross-hardware matrix: every golden-output gate must pass on Mac M2, Mac M3, and Windows/NVIDIA. Hardware inconsistency is the explicit verification axis, not an afterthought. + +## Status + +Hypothesis — internal dogfooding only. Not yet validated against external customers. Priority: P0 — ported alongside the render engine; cannot ship cross-hardware parity without it. diff --git a/docs/product/06-opportunities/03-edit-time-end-to-end.md b/docs/product/06-opportunities/03-edit-time-end-to-end.md new file mode 100644 index 0000000..6ac37a0 --- /dev/null +++ b/docs/product/06-opportunities/03-edit-time-end-to-end.md @@ -0,0 +1,47 @@ +# Opportunity: End-to-end edit time + +## Problem statement + +Producing one episode — long-form edit plus enough short-form clips for 1–2 weeks of promotion — costs a team member approximately 1.5 full days of hands-on work, making consistent biweekly publication unsustainable for a team with full-time jobs, and leaving no headroom to take on client work without sacrificing quality or burning out. + +## Evidence + +- Journey map stage: all stages (Sync → Transcribe → Edit → Render → short-form clip), but concentrated in Edit transcript and Render/export — no entries yet (pre-interviews). +- **Internal dogfooding (not yet externally validated):** + - "It takes a dedicated half a day at least to do a standard 30min–1hr podcast edit, considering we have to clean audio, sync audio and video, cut between speakers, add intro and name title cards, and outro." Source: [`00-problem-hypothesis.md`](../00-problem-hypothesis.md). + - "Creating short-form content through clipping for publicity from the long-form takes another day, especially when you want to create enough to distribute and promote across 1–2 weeks." Source: [`00-problem-hypothesis.md`](../00-problem-hypothesis.md). + - ragTech rotates editing across three people on a biweekly schedule — even with rotation, the process is described as "tedious" and trades off against audience engagement and product work. +- Lean Canvas §1 (Problem): both pain points above are listed as the top two problems verbatim. +- Lean Canvas §7 Cost Structure: `docs/FINANCIAL_PROJECTIONS.md` estimates hands-on time per tier (Essential ~1.0h, Growth ~2.0h, Premium ~4.0h) — these are the targets for a service business to be viable, and assume the pipeline handles the majority of mechanical work (sync, transcript-driven cuts, camera switching, short-form repurposing) without manual intervention per step. +- Lean Canvas §8 Key Metrics: "edit turnaround time per client project — raw footage to delivered edit" is a primary metric. + +## Proposed direction + +The pipeline should handle every mechanical, repeatable step without a human waiting on it: multi-angle sync, transcription, diarization, transcript-driven cut derivation, camera switching, intro/outro assembly, and short-form clip extraction from the same long-form source. The human's job is to review and correct the transcript doc and approve the cut — not to supervise each pipeline stage. Short-form repurposing should be a first-class output of the same pipeline run, not a separate manual pass with a different tool. + +This opportunity spans multiple pipeline stages and therefore multiple engineering epics — see related RFC context below for how they decompose. + +## Success metric + +- Full episode (long-form edit + 3–5 short-form clips) produced from raw footage in under 2 hours of calendar time, with under 30 minutes of human hands-on time (transcript review + approval). +- Lean Canvas §8 tie-in: "edit turnaround time per client project" hits same-day delivery for Essential/Growth tiers — raw footage in, delivered edit out, within a business day. +- Lean Canvas §7 unit economics: hands-on time per episode stays within the `docs/FINANCIAL_PROJECTIONS.md` estimates (Essential ~1.0h, Growth ~2.0h) so that the per-tier effective hourly rate (S$156–204/h) holds at the projected episode volume (Scenario B: ~17 episodes/month/person). + +## Related RFC / technical context + +This opportunity maps to multiple RFC build-order stages: + +| Pipeline stage | RFC reference | Build order | +|---|---|---| +| Sync (FFT cross-correlation) | RFC §Decision #4 — port from `AudioSyncer.js` faithfully, tuned constants preserved | Step 3 | +| Transcription | RFC §Decision #4 — `whisper-rs` direct binding, CUDA/Metal feature flags | Step 2 | +| Transcript editing (cut derivation, sentence merging) | RFC §Decision #4 — port from `edit-transcript.js`, preserve `PAUSE_THRESHOLD`, `WORD_DURATION_ESTIMATE`, `CUT_START_BIAS` exactly | Step 3 | +| Diarization, face detection, alignment, thumbnail removal | RFC §Decision #4 — keep as `tokio::process` subprocess calls indefinitely; these are third-party pretrained models | Step 3 | +| Short-form repurposing | Not yet scoped in RFC as a separate build step — currently `scripts/shorts/` in the existing codebase | TBD — depends on render engine (Step 1) being stable first | + +- [RFC §Context #4](../rfcs/0001-native-desktop-rewrite.md) — polyglot pipeline overview; JSON schema contracts (`transcript.json`, `camera-profiles.json`) as the stable interop boundary. +- RFC §Decision #3 — frozen JSON schemas as fixtures; new engine consumes them directly. + +## Status + +Hypothesis — internal dogfooding only. Not yet validated against external customers. Priority: P1 — depends on render speed (Opportunity #01) and hardware consistency (#02) being solved first; the pipeline's mechanical automation is only worth the investment if the render at the end isn't the bottleneck. diff --git a/docs/product/06-opportunities/04-transcript-timing-accuracy.md b/docs/product/06-opportunities/04-transcript-timing-accuracy.md new file mode 100644 index 0000000..848ba0a --- /dev/null +++ b/docs/product/06-opportunities/04-transcript-timing-accuracy.md @@ -0,0 +1,35 @@ +# Opportunity: Transcript token timing accuracy + +## Problem statement + +Transcript-driven cuts are only as precise as the word-level timestamps that define them. Current token timings are not accurate enough to make clean cuts without manual compensation — editors are forced to estimate and pad timings rather than trust the transcript, which defeats much of the automation benefit. + +## Evidence + +- **Internal dogfooding:** Word-level start/end timestamps (`t_dtw`, `t_end`) are not precise enough for frame-accurate cuts. The fallback when `t_end` is absent is a heuristic — `CUT_START_BIAS = 1.0` and `WORD_DURATION_ESTIMATE = 0.4s` in `scripts/edit-transcript.js` — and hook timing constants add further padding (`HOOK_TAIL_PAD_UNBOUNDED_SECONDS = 0.16s`, `HOOK_TAIL_PAD_BOUNDED_SECONDS = 0.02s` in `remotion/lib/hookTiming.ts`) specifically to absorb timing imprecision. These constants exist because timing cannot be trusted without them. +- **Internal dogfooding:** The current workaround for timing-sensitive cuts (e.g. hook clip boundaries) is to use `hookFrom?/hookTo?` bounds in the segment with manual time estimates — not visually placed, not frame-verified, just a number written into the doc. This requires the editor to estimate timing rather than observe it, making the process error-prone and non-repeatable. +- Root `CLAUDE.md` schema note: "`token.t_end` is populated only after forced alignment. Without it, `deriveCuts` falls back to `CUT_START_BIAS` heuristic." — the heuristic path is the default, not the exception. + +## Proposed direction + +Two improvements, in priority order: + +1. **Better forced alignment.** WhisperX forced alignment (currently a Python subprocess) produces `t_dtw` and `t_end` per token. The quality of these timestamps is the ceiling for cut precision — improving alignment accuracy directly improves every downstream cut without any other change. In the new codebase, `whisper-rs` for transcription and continued WhisperX subprocess for alignment is the planned path (RFC §Decision #4); the goal is that `t_end` is reliably populated and accurate enough that the `CUT_START_BIAS` fallback is never needed in practice. + +2. **Frame-accurate visual cut placement.** When alignment accuracy still isn't sufficient for a specific cut, the editor should be able to place the cut boundary visually (scrub to the frame, set the point) and have that propagate back into the transcript representation — not estimate a float and type it into the doc. This is the "synced with transcript" part and is addressed in [Opportunity 05 — visual transcript editor](./05-visual-transcript-editor.md). + +## Success metric + +- `t_end` populated for every token after the alignment stage — no token falls back to the `WORD_DURATION_ESTIMATE` heuristic in a normal pipeline run. +- Cut boundaries placed via transcript alone (no manual `hookFrom`/`hookTo` adjustment) are within 1 frame (≤16ms at 60fps) of the intended edit point, verified on real episode fixtures. +- The padding constants (`CUT_START_BIAS`, `HOOK_TAIL_PAD_UNBOUNDED_SECONDS`) can be reduced toward zero without introducing audible pops or visible frame bleed — meaning they're no longer load-bearing compensation for timing error. + +## Related RFC / technical context + +- [RFC §Decision #4](../rfcs/0001-native-desktop-rewrite.md) — Transcription: port to `whisper-rs` with CUDA/Metal feature flags. Forced alignment (WhisperX): kept as `tokio::process` subprocess call. +- RFC §Decision #4 — "Port the algorithm from `scripts/edit-transcript.js`... preserving the tuned constants (`PAUSE_THRESHOLD`, `WORD_DURATION_ESTIMATE`, `CUT_START_BIAS`) exactly — they encode real editorial behavior." This is true for the initial port; reducing these constants is a post-port accuracy goal, not a v1 requirement. +- Root `CLAUDE.md` — `WORD_DURATION_ESTIMATE`, `CUT_START_BIAS`, `HOOK_TAIL_PAD_UNBOUNDED_SECONDS`, `HOOK_TAIL_PAD_BOUNDED_SECONDS` — the full list of constants that exist to compensate for timing imprecision. + +## Status + +Validated — internal dogfooding. Pain is directly observed and the workaround (manual time estimation for hook boundaries) is actively in use. Priority: P1 — depends on the transcription pipeline being ported (RFC Build Order Step 2) before this can be improved. diff --git a/docs/product/06-opportunities/05-visual-transcript-editor.md b/docs/product/06-opportunities/05-visual-transcript-editor.md new file mode 100644 index 0000000..f560883 --- /dev/null +++ b/docs/product/06-opportunities/05-visual-transcript-editor.md @@ -0,0 +1,42 @@ +# Opportunity: Visual transcript editor + +## Problem statement + +The current editing interface is a plain text file (`transcript.doc.txt`) with a custom cue syntax edited in a code editor. This is not intuitive for timing-sensitive or visual decisions — the editor cannot see the video while placing cuts, cannot verify a cut boundary without rendering, and must write timing estimates as raw numbers rather than placing them on a visual timeline. The gap between "editing the text" and "seeing the result" is the dominant source of trial-and-error in the current workflow. + +## Evidence + +- **Internal dogfooding:** Transcript editing currently happens in a VSCode extension (`vscode-transcript-language/src/extension.js`) that provides syntax highlighting for the cue format but no video preview, no waveform, and no playback. Cut decisions are made by reading the text, not by watching the moment. +- **Internal dogfooding:** Hook clip boundaries (`hookFrom?/hookTo?`) require manually typing float timestamps estimated from memory — not from scrubbing to the frame. This is the direct consequence of having no visual editor synced with the transcript; see [Opportunity 04 — timing accuracy](./04-transcript-timing-accuracy.md) for the alignment-side complement to this problem. +- **Internal dogfooding:** The cue syntax itself (curly braces for word cuts, `> SPEAKER` splits, `> CAM` directives, `> HOOK` annotations) is expressive but not discoverable. New team members need to read the format documentation to use it; there's no visual affordance showing which words are cut, which segments are camera-switched, or where hooks begin and end. +- Root `CLAUDE.md` Phase 8 plan: `app/editor/page.tsx` (transcript editor) and `app/editor/Timeline.tsx` (630-line timeline component) are already identified as requiring a `PreviewPlayer`, scroll-sync, waveform, and decomposition — the existing Next.js codebase already reached this conclusion, it just hasn't been built yet. + +## Proposed direction + +A visual transcript editor where the text representation and the video/timeline representation stay in sync — editing either one updates the other. Specifically: + +- **Waveform + playhead** synced to the transcript text: clicking a word in the transcript scrubs the video to that word's timestamp; moving the playhead in the timeline highlights the corresponding word in the transcript. +- **Visual cut placement:** dragging a cut boundary on the timeline updates `t_dtw`/`t_end` in the underlying transcript; the text representation reflects the cut without requiring the editor to type a float. +- **In-context preview:** the edited cut plays back immediately in the editor without a full render — the same compositor used for final render (RFC Decision §2) drives the preview, so what you see in the editor is what renders. +- The cue syntax (curly braces, directives) becomes the *persistence format*, not the *editing interface* — power users can still edit the text directly, but the visual layer is the primary interaction for timing-sensitive decisions. + +This is one of the motivations for the egui + wgpu native GUI in RFC Decision §2, which explicitly calls out the timeline/waveform editor as a natural fit for egui's immediate-mode model and notes the paradigm carries over from the existing Next.js canvas timeline. + +## Success metric + +- An editor can place a cut boundary to within 1 frame (≤16ms at 60fps) by scrubbing visually — no float estimation required. +- An editor can complete a hook clip boundary (`hookFrom`/`hookTo`) by setting in/out points on the timeline rather than writing numbers in the doc. +- The in-editor preview matches the final render output for the same segment — no "looked fine in preview, broken in render" class of bug. +- A new team member can make a basic cut (mark a word range, preview it, approve it) without reading the cue syntax documentation. + +## Related RFC / technical context + +- [RFC §Decision #2](../rfcs/0001-native-desktop-rewrite.md) — Native GUI (egui + wgpu); one compositor library shared by preview surface and final-render encoder. The "preview matches render" guarantee is structural, not incidental. +- RFC §Decision #2 — "egui's immediate-mode painting model is a good fit for a timeline/waveform editor specifically, since per-frame redraw of playheads, waveforms, and clip lanes is the natural way that kind of UI already worked in the existing custom-canvas Next.js timeline editor." +- RFC §Build Order Step 4 — Native GUI built after the compositor API is stable (Step 1), so the editor isn't chasing a moving target. +- Root `CLAUDE.md` Phase 8 — `app/editor/page.tsx`, `app/editor/Timeline.tsx` — existing Next.js editor identified for PreviewPlayer + scroll-sync + waveform; this opportunity supersedes that plan in the new codebase. +- Root `CLAUDE.md` — `vscode-transcript-language/src/extension.js` — the current editing surface being replaced. + +## Status + +Validated — internal dogfooding. The pain (blind text-based cuts, float estimation for hook timings, no in-context preview) is directly observed in every episode edit. Priority: P2 — depends on the render/compositor engine (RFC Build Order Step 1) being stable before the GUI is built on top of it. diff --git a/docs/product/06-opportunities/06-multi-brand-support.md b/docs/product/06-opportunities/06-multi-brand-support.md new file mode 100644 index 0000000..97a617e --- /dev/null +++ b/docs/product/06-opportunities/06-multi-brand-support.md @@ -0,0 +1,44 @@ +# Opportunity: Multi-brand support + +## Problem statement + +The current tool is built for exactly one brand — ragTech — with its logo, colors, Nunito font, Techybara mascot, overlays, intro/outro music, and host identities baked into the codebase. Editing a client's podcast means their footage runs through ragTech's visual identity. The service business (studio-rental partnership, any future direct client work) cannot operate without brand-per-job support; every client has their own logo, palette, font, and show format. + +## Evidence + +- **Internal dogfooding:** Brand assets are hardcoded throughout the current codebase — `public/brand.json` holds a single brand config, `public/assets/` contains only ragTech team images and Techybara PNGs, intro/outro music (`public/sounds/`) is ragTech-specific, and `remotion/lib/brandRegistry.ts` implements only the ragTech brand. `OverlayRenderer.tsx` dispatches via `CORE_TEMPLATE_MAP + getBrandOverlays(brand.id)` but `getBrandOverlays` only returns ragTech overlays. Source: root `CLAUDE.md`. +- **Internal dogfooding:** The in-place refactor to fix this (Phase 0.5, `refactor/p0-brand` branch) is already underway in the current codebase — `public/brand.json` is being migrated to `brands/ragtech/brand.json`, brand overlays are being moved to `brands/ragtech/components/`, and `brandRegistry.ts` is being updated to load brands dynamically. The fact that a dedicated refactor phase exists confirms the hardcoding was a real constraint, not a hypothetical one. +- **Service business requirement:** Lean Canvas §2 Segment 1 (podcast studio-rental partnership, active deal in motion) routes client editing work to us as a fulfillment partner. Those clients have their own podcast brands — their own logos, color palettes, fonts, hosts, overlays, and show music. Delivering their edited footage with ragTech's branding would not be deliverable. +- **Problem hypothesis:** "Now we want to scale up to editing for other podcasts so this can be a viable business." Multi-brand is the prerequisite, not a nice-to-have. + +## Proposed direction + +Brand configuration should be a first-class, per-job input to the pipeline — not a compile-time or repo-level constant. Each brand specifies: + +- **Identity:** name, logo (with transparent background), primary/secondary/accent colors, font(s). +- **Hosts:** per-host name, image, camera angle mapping — equivalent to the current `speakers` section of `camera-profiles.json`, but owned by the brand config. +- **Audio:** intro/outro music file, background music file (or none). +- **Overlays and templates:** which intro sequence, which name card style, which lower-thirds, which outro — either picking from a library of core templates or supplying brand-specific overlay components. +- **Mascot/visual assets:** optional; ragTech has Techybara, another brand may have nothing or something different. + +The compositor selects brand assets at job time, not at build time. Adding a new brand means adding a new brand directory and config — no code change to the core pipeline. + +The new codebase should build this in from the start. The existing codebase is retrofitting it via Phase 0.5; the rewrite should not repeat that pattern. + +## Success metric + +- A pipeline run for Brand A and Brand B on the same raw footage produces two different outputs — correct logos, colors, fonts, host name cards, intro/outro — with zero manual file-swapping between runs. +- Adding a third brand requires only: a new brand config file + asset directory. No changes to pipeline code, compositor, or overlay logic. +- ragTech's own episodes continue to render correctly as Brand 0 — the first brand in the system, not a special case. + +## Related RFC / technical context + +- Root `CLAUDE.md` Phase 0.5 (`refactor/p0-brand`) — the in-place brand abstraction work underway in the current codebase. The new codebase should treat this as a solved design problem (the Phase 0.5 schema is the reference), not repeat the discovery. +- Root `CLAUDE.md` — `remotion/types/brand.ts`: "Brand design tokens + extended identity/hosts/mascot/audio; `id: string` field required by registry." This type shape is the starting point for the new codebase's brand config schema. +- Root `CLAUDE.md` — `remotion/lib/brandRegistry.ts`: `getBrandOverlays(brandId)` registry pattern — the right abstraction, currently only ragTech-populated. +- Root `CLAUDE.md` — `OverlayRenderer.tsx`: "Remove remaining brand hardcoding (Phase 0.5 Steps 6–7)" — confirms hardcoding is still present in the current codebase. +- [RFC §Decision #3](../rfcs/0001-native-desktop-rewrite.md) — frozen JSON schemas as the interop contract. Brand config should be a versioned schema alongside `transcript.json` and `camera-profiles.json`, not ad hoc. + +## Status + +Validated — internal dogfooding + active service business requirement. The constraint is directly observed (ragTech branding is hardcoded), and the studio-rental partnership deal makes multi-brand support a prerequisite for the service business to function. Priority: P1 — required before taking on any client work outside ragTech's own episodes. diff --git a/docs/product/README.md b/docs/product/README.md index 97c5bcf..f5d1e7b 100644 --- a/docs/product/README.md +++ b/docs/product/README.md @@ -26,9 +26,23 @@ Do these roughly in sequence — each one narrows/grounds the next. Don't skip s | 6 | [`05-personas.md`](05-personas.md) | 1–2 lightweight personas distilled from interview clusters | after interviews | | 7 | [`06-opportunities/`](06-opportunities/) | One doc per opportunity surfaced above — the direct bridge to epics/issues in the new repo | ongoing | -## What "done" looks like before writing the first engineering issue in the new repo +## What "done" looks like at each phase -- Lean Canvas has been revised at least once based on real interview feedback (not just your own assumptions). -- At least 5 interviews with people outside ragTech, synthesized (not just raw transcripts sitting unread). -- A journey map that names specific pain points with evidence (a quote, a timing observation), not just "editing is slow." -- Each opportunity doc traces back to a specific journey-map pain point or interview finding — if you can't point to the evidence, it's not ready to become an epic yet. +### Phase 1 — internal dogfooding (current phase) + +The tool stays internal; what we sell externally is the editing service, not the app. Engineering issues can begin once: + +- Opportunity docs exist in `06-opportunities/` for each area of work, with evidence explicitly labeled (internal dogfooding, RFC analysis, or financial modeling — not just "we find this annoying"). +- Each opportunity doc links to the RFC section or product doc that grounds it — if that link doesn't exist, it's a hypothesis without a paper trail, not a ready opportunity. +- The journey map (`04-journey-map.md`) is filled in from ragTech's own pipeline experience, with pain points tied to observed time costs or named failure modes. + +Interviews and external validation are **not** a gate for Phase 1. The studio-rental partnership (Lean Canvas §2, Segment 1) is the one near-term external relationship that matters; it may produce real-usage feedback, but it does not replace the internal dogfooding track. + +### Phase 2 — before scaling externally / self-serve product + +When we are satisfied with the product internally (multiple videos, multiple brands, stable quality) and want to begin serving external customers directly — not just fulfilling via the studio partner — the following should be true before treating any new opportunity as validated: + +- At least 5 interviews with people outside ragTech, synthesized in `03-customer-interviews/synthesis.md`. +- The journey map revised at least once based on real interview findings — not just ragTech's own pipeline. +- Lean Canvas revised at least once based on real interview feedback. +- Each new opportunity doc cites a specific journey-map pain point or interview quote as evidence — if it can't, it stays labeled "hypothesis." From 4803b6fe7a583f00c23f03aa4e56a9c9393d076a Mon Sep 17 00:00:00 2001 From: Natasha Ann Date: Sun, 12 Jul 2026 23:13:53 +0800 Subject: [PATCH 5/7] docs(product): add journey map --- docs/product/04-journey-map.md | 38 ++++++++++++++++++++-------------- 1 file changed, 23 insertions(+), 15 deletions(-) diff --git a/docs/product/04-journey-map.md b/docs/product/04-journey-map.md index a831046..d3c5be3 100644 --- a/docs/product/04-journey-map.md +++ b/docs/product/04-journey-map.md @@ -6,33 +6,41 @@ Map one representative persona/segment at a time — if you have multiple segmen ## Persona / segment for this map -*Which persona from `05-personas.md` (or interview segment) this map represents.* +**ragTech Operator** — engineer-editor rotating through the pipeline biweekly. See [`05-personas.md` Persona 1](05-personas.md). -- +*A second map for an external client operator (someone outside the ragTech team using the tool to edit their own podcast) should be created once the studio-rental partnership produces real usage — do not fabricate it from assumptions.* ## Stages -*Adapt these to what interviews actually revealed — this is a reasonable starting skeleton based on ragTech's own pipeline, but a target customer's real workflow may have different or extra stages (e.g. scripting/planning before recording, client review/approval after editing).* +*Stages reflect ragTech's actual pipeline as of 2026-07. Short-form clip creation is a separate stage (not part of the scaffolded template) because it is a full additional pipeline pass, not a publish-side step. Evidence source for all rows: internal dogfooding unless noted otherwise.* | Stage | What they're doing | What they're thinking/feeling | Pain points (with evidence) | Time cost | Opportunity | |-------|----------------------|----------------------------------|-------------------------------|-----------|--------------| -| Record | | | | | | -| Sync (multi-angle/audio) | | | | | | -| Transcribe | | | | | | -| Diarize/label speakers | | | | | | -| Edit transcript (cuts, corrections) | | | | | | -| Camera/framing decisions | | | | | | -| Render/export | | | | | | -| Publish/distribute | | | | | | +| Record | Recording the episode across 3 camera angles + separate audio source | "Did the cameras all start at roughly the same time? Was audio clean?" | Sync issues are only discovered post-recording — no way to know during the session if an angle will be unusable | Low; not a pipeline step we control | — | +| Sync | Running FFT cross-correlation (`scripts/sync/AudioSyncer.js`) to align camera angles to audio; verifying `synced-output-{N}.mp4` files by spot-checking | "Did it actually sync, or is there a drift I'll only notice mid-edit?" | No automated quality signal for sync accuracy — verification is manual spot-checking; silent failure (drift present but not obvious until a cut looks wrong) produces errors that are expensive to debug later. Hardware inconsistency: `AudioSyncer.js` branches on platform for encoder but downstream scripts hardcode `libx264` regardless. | ~15–30 min automated + manual spot-check | [Opportunity 02](06-opportunities/02-hardware-inconsistency.md) | +| Transcribe | Running whisper.cpp via subprocess; waiting for transcript.raw.json | "How many timing errors will I have to correct this time?" | Token start/end timestamps (`t_dtw`, `t_end`) are not frame-accurate; `t_end` only populated after forced alignment, so `deriveCuts` falls back to `CUT_START_BIAS = 1.0` / `WORD_DURATION_ESTIMATE = 0.4s` heuristics as default. Accuracy determines how much downstream manual correction is needed. | ~10–30 min depending on episode length and hardware | [Opportunity 04](06-opportunities/04-transcript-timing-accuracy.md) | +| Diarize / label speakers | Running pyannote diarization + WhisperX forced alignment (Python subprocesses); manually assigning speaker names to diarized segments | "Did it split the speaker turns correctly? How many do I need to fix?" | Diarization errors (missed turns, wrong boundaries) require manual correction in the doc; no visual way to verify speaker labels against the video. Forced alignment populates `t_end` but quality is uneven — some tokens still land on wrong frames. | ~20–40 min to verify and correct | [Opportunity 04](06-opportunities/04-transcript-timing-accuracy.md) | +| Edit transcript | Opening `transcript.doc.txt` in VSCode; marking word cuts with `{curly braces}`, adding `> SPEAKER` splits, `> CAM` directives, `> HOOK` annotations; running `merge-doc` to apply edits | "I can't see the video while I'm doing this. I'm guessing whether this cut will sound right." | (1) No video preview while editing — all cut decisions are made by reading text, not watching the moment. (2) Hook clip boundaries (`hookFrom`/`hookTo`) require typing estimated float timestamps, not scrubbing to a frame. (3) Cue syntax is not discoverable — new team members need documentation. (4) No feedback on cut quality until a full render completes. (5) Single-brand only — ragTech assets are hardcoded; client footage cannot be edited with a different brand. | Largest hands-on time block; majority of the "half a day minimum" per `00-problem-hypothesis.md` | [Opportunity 04](06-opportunities/04-transcript-timing-accuracy.md), [Opportunity 05](06-opportunities/05-visual-transcript-editor.md), [Opportunity 06](06-opportunities/06-multi-brand-support.md) | +| Camera / framing setup | Running `setup-camera.js` (face detection via mediapipe); reviewing and adjusting closeup viewports per speaker in the camera GUI (`app/camera/page.tsx`) | "Did face detection crop correctly? I need to check every angle manually." | Face detection requires manual verification for every episode — viewport crops drift when speakers move. Adjustments cannot be previewed against real footage in context; changes are applied and then checked by re-running. | ~30–60 min | [Opportunity 05](06-opportunities/05-visual-transcript-editor.md) | +| Render / export | Triggering Remotion render (headless Chromium); waiting; reviewing output | "I've started the render. I can't use this machine for the next several hours. I hope it doesn't fail." | (1) 6–18 hours for a 60-minute 3-angle episode at 60fps — same-day review is impossible. (2) Silent hardware degradation: Windows/NVIDIA silently falls back to `libx264`; no error, just slower encode and inconsistent output. (3) If output has a bug, fix + re-render adds another 6–18 hours. | 6–18 hours elapsed; ~10 min hands-on | [Opportunity 01](06-opportunities/01-render-speed.md), [Opportunity 02](06-opportunities/02-hardware-inconsistency.md) | +| Short-form clip creation | Separate pipeline pass: running `scripts/shorts/wizard`, selecting clip ranges from long-form transcript, running sync/transcribe/align/edit/render again for portrait format | "This is a whole separate day of work to get 3–5 clips for the week's promotion." | Short-form is not a first-class output of the long-form pipeline — it is a full second pipeline pass with its own transcription, editing, and render cycle. The same render bottleneck (Opportunity 01) applies again. Requires another ~1 full day per `00-problem-hypothesis.md`. | ~1 full additional day | [Opportunity 01](06-opportunities/01-render-speed.md), [Opportunity 03](06-opportunities/03-edit-time-end-to-end.md) | +| Publish / distribute | Uploading long-form to YouTube/Spotify/Apple; short-form to Instagram/TikTok/LinkedIn | "Just need to get this uploaded before the week is out." | Not a pipeline pain point we're solving — out of scope for v1. | Variable | — | ## Biggest pain point overall -*The single stage/moment that came up most often or most intensely across interviews. This is usually where the first opportunity doc should focus.* +**Render time (Opportunity 01) and the transcript editing loop (Opportunities 04 + 05) together.** Render time is the single largest calendar-time cost and the hard blocker on same-day delivery. The transcript editing loop (inaccurate timing → blind text editing → full render to verify) is the largest *hands-on* time cost and the most friction-per-minute of work. They compound: a mis-timed cut discovered post-render means correcting the transcript and waiting another 6–18 hours. -- +Evidence source: internal dogfooding — directly observed across multiple episode production cycles. Not yet validated externally. ## Evidence trail -*For each pain point above, link back to the specific interview # and quote in `03-customer-interviews/synthesis.md` that supports it. If a row has no evidence trail, treat it as your own assumption, not a validated finding — mark it as such.* +All rows: **internal dogfooding** — direct observation from building and operating the ragTech pipeline. No external interview evidence yet (Phase 2, per `README.md`). -- +Specific supporting sources per row: +- Sync: `scripts/sync/AudioSyncer.js` (platform branching); RFC §Context #3 (hardcoded `libx264` in downstream scripts) +- Transcribe: root `CLAUDE.md` schema note on `t_end`; `WORD_DURATION_ESTIMATE`, `CUT_START_BIAS` constants in `scripts/edit-transcript.js` +- Diarize: root `CLAUDE.md` pipeline diagram; WhisperX forced alignment subprocess +- Edit transcript: direct operator experience; `00-problem-hypothesis.md` "half a day minimum"; `vscode-transcript-language/` as current editing surface +- Camera setup: `app/camera/page.tsx`; `scripts/camera/setup-camera.js` +- Render/export: RFC §Context #1 (2–5fps ceiling, 6–18h for 60-min episode); RFC §Context #2 (Windows/NVIDIA silent fallback) +- Short-form: `00-problem-hypothesis.md` "another full day"; `scripts/shorts/wizard` From 36277574e9cde0a9c386f65fa3549963bc68ba34 Mon Sep 17 00:00:00 2001 From: Natasha Ann Date: Fri, 17 Jul 2026 22:14:32 +0800 Subject: [PATCH 6/7] docs: add epic doc --- docs/product/02-stakeholder-map.md | 30 +++-- docs/product/07-epics.md | 186 +++++++++++++++++++++++++++++ 2 files changed, 203 insertions(+), 13 deletions(-) create mode 100644 docs/product/07-epics.md diff --git a/docs/product/02-stakeholder-map.md b/docs/product/02-stakeholder-map.md index ae70041..4d8c888 100644 --- a/docs/product/02-stakeholder-map.md +++ b/docs/product/02-stakeholder-map.md @@ -4,32 +4,36 @@ Quick exercise — mostly internal right now. Purpose: know who has decision rig ## Core team -*Confirm roles and decision rights explicitly — don't assume everyone agrees who owns what.* +*Decision rights and primary concerns listed below are a starting point — confirm as a team and update if they don't match how you actually operate.* | Name | Role | Decision rights (what can they say yes/no to alone?) | Primary concern | |------|------|--------------------------------------------------------|------------------| -| Natasha | Software Engineer | | | -| Saloni | Software Developer | | | -| Victoria | Solutions Engineer | | | +| Natasha | Software Engineer | Technical architecture, RFC decisions, codebase structure, engineering issue scope | Reliability and correctness of the new codebase; cross-hardware consistency | +| Saloni | Software Developer | Feature implementation decisions, script/pipeline logic | Pipeline accuracy (transcription, diarization, sync); editing quality | +| Victoria | Solutions Engineer | External-facing scope, client relationship/partnership decisions, pricing and packaging | Deliverability to clients; turnaround time; output quality for service business | + +All three share decision rights over: product scope, which features are in or out of the current phase, and whether to take on new client work. ## Extended / occasional stakeholders -*Advisors, investors, contractors, anyone with influence but not day-to-day involvement. If none yet, write "none yet" — don't leave it blank, since that's a different fact than "not applicable."* +*Advisors, investors, contractors, anyone with influence but not day-to-day involvement.* -| Name/Group | Relationship | Influence level | What they need to be kept informed of | -|------------|---------------|------------------|------------------------------------------| -| | | | | +None yet. ## Prospective external stakeholders (product path) -*People not yet involved but relevant if this becomes a product — early design partners you already know, a target customer community, potential early beta users.* +*People not yet involved but relevant as the service business and eventually the product develop.* | Name/Group | Relationship | Why they matter | Next step to engage them | |------------|---------------|------------------|----------------------------| -| | | | | +| Podcast studio rental company (unnamed) | Active deal in motion — routes client editing work to ragTech as a fulfillment partner | First real external distribution channel; their clients are the first non-ragTech footage the tool will need to handle (different brands, different formats). See Lean Canvas §2 Segment 1. | Convert to a design partner relationship — agree on what a delivered edit looks like for their first client, use that as the acceptance criteria for multi-brand support (Opportunity 06) | +| Podcast studio clients (via partner) | Indirect — ragTech never directly interacts with them; the studio owns that relationship | Their footage and brand requirements are the first real test of multi-brand support and non-ragTech pipeline inputs | No direct engagement needed in Phase 1 — fulfill via the studio partner; gather indirect feedback through them | -## Interest/influence grid (optional, quick) +## Interest/influence grid -*Plot each stakeholder above: high-influence/high-interest people need active management; high-influence/low-interest need to be kept satisfied without over-communicating; low-influence/high-interest are good candidates for early interviews/beta.* +| | High influence | Low influence | +|---|---|---| +| **High interest** | Natasha, Saloni, Victoria (core team) | Future self-serve customers (Phase 2+) | +| **Low interest** | Studio rental partner (cares about delivery quality and SLA, not how the tool works) | — | -- +Active management: core team only. Studio partner: keep informed on turnaround time and output quality; do not expose internal tooling complexity. diff --git a/docs/product/07-epics.md b/docs/product/07-epics.md new file mode 100644 index 0000000..728f2f0 --- /dev/null +++ b/docs/product/07-epics.md @@ -0,0 +1,186 @@ +# Epic Map + +This document is the direct bridge from product discovery to engineering. Each epic here maps to one or more opportunity docs in `06-opportunities/`, is sequenced against the RFC build order, and is broken into features at enough resolution to write a GitHub issue with acceptance criteria from each one. + +**How to use this in the new repo:** +- Each epic becomes a GitHub Epic (or milestone, depending on how the new repo structures work). +- Each feature under an epic becomes one or more GitHub issues, with acceptance criteria derived from the success metrics in the linked opportunity doc. +- Epic 0 must be the first epic completed — its outputs (test fixtures, golden-frame baselines, AI harness) are the acceptance criteria infrastructure that every subsequent epic depends on. +- The spike must be resolved before Epic 1 issues are written — its outcome determines which FFmpeg binding the entire render engine is built against. It can run in parallel with Epic 0 since it lives in the current repo. + +--- + +## Pre-epic: FFmpeg vs gstreamer-rs spike + +**Not an epic — a one-off technical investigation that must complete before Epic 1 begins.** + +The RFC calls this out explicitly as non-blocking to the plan but blocking to the render engine epic: the decode/encode binding choice determines every subsequent API call in the compositor. Deciding mid-epic would force a rewrite of already-written code. + +**Question to answer:** Is `ffmpeg-next` (direct FFmpeg C API bindings) or `gstreamer-rs` the better foundation for the render engine on the target hardware matrix (Mac M2/M3 VideoToolbox, Windows NVIDIA NVENC/NVDEC)? + +**Evaluation criteria:** +- Hardware codec access: can it reach VideoToolbox on Mac and NVENC/NVDEC on Windows without extra shims? +- Seek/decode accuracy: frame-accurate seek to an arbitrary timestamp (required for jump-cut editing). +- Build complexity: does it cross-compile cleanly on all three target platforms from CI? +- Maintenance: crate maturity, update frequency, known soundness issues. + +**Spike output:** a short findings doc (not a polished RFC — a decision record with the test results) committed to the spike branch. The finding becomes the `encoderProfile` design decision that Epic 1 is built against. + +**Where to run it:** a `spikes/rust-decode-spike/` directory in the current repo (deckcreate), on its own branch. Findings port to the new repo; the spike code does not. + +--- + +## Epic 0: Project Setup and AI Harness + +**Opportunities addressed:** None directly — this epic produces the infrastructure that makes all other epics verifiable and agent-driveable. + +**RFC reference:** RFC §Open Questions — repository location, Agent Implementation Convention. RFC §Verification §6 — golden-output gate per stage; cross-hardware matrix; schema conformance tests. All of these require the fixtures, baselines, and tooling this epic creates. + +**Problem it solves:** Without test fixtures and golden-frame baselines, Epic 1's acceptance criteria ("output passes PSNR gate vs Remotion baseline") have nothing to diff against — you can't close Epic 1 as done. Without a CLAUDE.md and implementation convention, coding agents working in the new repo have no shared understanding of how work is structured, how to validate a change, or what "done" means per issue. This epic makes all subsequent epics executable and verifiable. + +**Definition of done:** A coding agent can be handed any Epic 1 feature issue, find the CLAUDE.md, run `cargo test`, confirm the golden-frame baseline exists for that fixture, and know exactly what a passing acceptance check looks like — without reading anything outside the new repo. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F0.1 | Cargo workspace structure | Multi-crate Cargo workspace with one crate per major subsystem: `crates/compositor`, `crates/transcription`, `crates/pipeline`, `crates/gui`, `crates/brands`; shared types crate (`crates/types`) for `transcript.json` and `camera-profiles.json` structs | `cargo build --workspace` succeeds on Mac and Windows with no warnings; crate boundaries match the Epic 1–5 feature split so no epic needs to reach across the wrong crate | +| F0.2 | CLAUDE.md — AI harness | Agent implementation convention for the new repo: crate layout, how to run tests (`cargo test`, `cargo clippy`, `cargo fmt --check`), implementation doc format (one doc per multi-step task in `docs/implementation-guides/`), status check per step, one commit per step, per-PR checklist adapted for Rust | An agent given a feature issue can orient, find the right crate, run the test suite, and know what done looks like without any out-of-repo context | +| F0.3 | CI/CD cross-platform matrix | GitHub Actions workflow: build + test on Mac ARM (M-series), Windows x64 with NVIDIA GPU runner; clippy and rustfmt checks on every PR; test matrix fails fast if any target fails | Green CI required to merge any PR; a failure on Windows/NVIDIA is as blocking as a failure on Mac | +| F0.4 | Test fixture export | Export real `transcript.json` + `camera-profiles.json` from at least two already-shipped episodes from the current production pipeline; commit as versioned fixtures in `tests/fixtures/` | Fixtures round-trip through the new repo's Rust schema structs without loss; used as the canonical input for all Epic 1–3 acceptance tests | +| F0.5 | Golden-frame baseline export | For each fixture episode, render a reference set of frames using the current Remotion pipeline and commit as baselines in `tests/baselines/`; include a PSNR threshold and audio hash per fixture | Epic 1's compositor output can be diffed against these baselines with a single test command; baselines are reproducible from the current pipeline if they need to be regenerated | +| F0.6 | Shared test utilities | Common test helpers in `crates/test-utils`: fixture loader, PSNR calculator, audio hash utility, frame extractor — consumed by all epic test suites | Any crate's test suite can load a fixture and run a PSNR diff in under 5 lines of test code | +| F0.7 | Dev tooling | `rustfmt.toml`, `.clippy.toml`, pre-commit hooks (or equivalent) enforcing format and lint on changed files; `Makefile` or `justfile` with canonical commands (`make test`, `make lint`, `make baseline`) | A new contributor (or agent) can set up the dev environment and run the full check suite with documented commands; no manual tool invocation required | + +--- + +## Epic 1: Native Render Engine + +**Opportunities addressed:** [Opportunity 01 — Render speed](06-opportunities/01-render-speed.md), [Opportunity 02 — Hardware-consistent output](06-opportunities/02-hardware-inconsistency.md) + +**RFC reference:** Build Order Step 1. Must be the first epic completed — all other epics depend on a stable compositor API. + +**Problem it solves:** The current Remotion/headless-Chromium render path has a hard 2–5fps ceiling with no GPU path, producing 6–18 hour renders for a 60-minute episode. Hardware inconsistency (Windows/NVIDIA silently falling back to software encode) means output quality depends on which machine ran the job. + +**Definition of done:** A render of the same episode fixture on Mac M2, Mac M3, and Windows/NVIDIA produces output that passes a PSNR/hash diff gate and completes in under 1 hour — verified across all three hardware targets. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F1.0 | Spike outcome applied | FFmpeg binding (or gstreamer-rs) chosen from the spike; dependency declared in `Cargo.toml` | Spike findings doc exists and is linked from this epic's first issue | +| F1.1 | Multi-angle compositor | wgpu-based compositor that stacks multiple decoded video planes, applies viewport crops, and composites to an output surface — same code path for preview and final render | Composites 3 angles with correct viewport transforms matching the `camera-profiles.json` spec; verified against golden frames from the current Remotion output | +| F1.2 | VideoToolbox decode/encode path | Hardware-accelerated H.264 decode and encode via AVFoundation + VideoToolbox on macOS | Render on Mac M2 and Mac M3 completes in under 1 hour for a 60-min episode; output passes PSNR gate vs Remotion baseline | +| F1.3 | NVENC/NVDEC decode/encode path | Hardware-accelerated H.264 decode (NVDEC) and encode (NVENC) on Windows/NVIDIA | Render on Windows/NVIDIA completes in under 1 hour; output passes PSNR gate; no silent fallback to libx264 | +| F1.4 | Real hardware capability probing | Replaces the `process.platform`/`process.arch` string heuristic in `hardware.ts` with actual GPU/codec capability queries at startup | Correctly identifies VideoToolbox on Mac and NVENC on Windows/NVIDIA; logs the resolved encoder profile at startup; does not silently fall back | +| F1.5 | Unified encoder profile mapping | Single `encoderProfile → FFmpeg/codec args` mapping consumed by every pipeline stage — no per-script encoder decisions | No pipeline script sets its own encoder; all route through the profile mapping; adding a new profile requires one code change in one file | +| F1.6 | Golden-frame validation suite | Automated diff of new-engine output against Remotion output for real-episode fixtures — PSNR/hash diff for video, audio-hash diff for audio | Suite runs on CI across Mac M2, Mac M3, and Windows/NVIDIA; fails if PSNR drops below threshold or audio hash mismatches | +| F1.7 | `transcript.json` + `camera-profiles.json` schema ingestion | Rust structs for both schemas; round-trips real fixture files without loss | Schema conformance tests pass on all fixture files exported from the current production pipeline | + +--- + +## Epic 2: Transcription and Alignment + +**Opportunity addressed:** [Opportunity 04 — Transcript token timing accuracy](06-opportunities/04-transcript-timing-accuracy.md) + +**RFC reference:** Build Order Step 2. Isolated and low-risk; can begin once Epic 1's compositor API is stable enough that the transcription output format is known. + +**Problem it solves:** Token timestamps are not frame-accurate — `t_end` is only populated after forced alignment, and `deriveCuts` falls back to `CUT_START_BIAS`/`WORD_DURATION_ESTIMATE` heuristics as the default path. Inaccurate timestamps mean every cut decision has a margin of error that compounds downstream. + +**Definition of done:** `t_end` is reliably populated for every token after alignment; cut boundaries derived from the transcript alone land within 1 frame (≤16ms at 60fps) of the intended edit point on real episode fixtures. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F2.1 | whisper-rs transcription | Direct `whisper-rs` binding replacing the current subprocess + VTT-parsing contract; CUDA and Metal feature flags enabled | Transcription output matches current Whisper output on real episode fixtures; runs on Mac (Metal) and Windows (CUDA) without subprocess | +| F2.2 | WhisperX forced alignment subprocess | `tokio::process` call to WhisperX; parses output into `t_dtw`/`t_end` per token | `t_end` populated for every token in alignment output; round-trips to the `transcript.json` schema | +| F2.3 | Alignment accuracy validation | Automated test comparing aligned token boundaries against human-verified ground truth on a sample episode | No token's `t_end` misses its true boundary by more than 1 frame on the validation sample; `WORD_DURATION_ESTIMATE` fallback is not invoked in a normal run | + +--- + +## Epic 3: Pipeline Orchestration + +**Opportunity addressed:** [Opportunity 03 — End-to-end edit time](06-opportunities/03-edit-time-end-to-end.md) + +**RFC reference:** Build Order Step 3. Depends on Epics 1 and 2 being stable — the orchestration wires around the proven render core. + +**Problem it solves:** Every mechanical, repeatable step (sync, transcribe, diarize, cut derivation, camera setup, short-form extraction) currently requires manual supervision and produces a full additional day of work for short-form clips. The pipeline should handle all of it with the human's role limited to transcript review and approval. + +**Definition of done:** Raw footage in → long-form edit + 3–5 short-form clips out, with under 30 minutes of human hands-on time (transcript review + approval). Measured on a real episode. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F3.1 | Audio/video sync | Port of FFT cross-correlation sync from `scripts/sync/AudioSyncer.js` into Rust; deterministic tie-break and SNR reliability check preserved; tuned constants carried over exactly | Sync output matches current AudioSyncer output on real multi-angle fixtures; offset measured within ±1 frame | +| F3.2 | Diarization subprocess | `tokio::process` call to pyannote; output parsed into speaker-turn segments | Speaker turns correctly segmented on real episode fixtures; speaker labels assignable from output | +| F3.3 | Speaker assignment | Maps diarization output to speaker names using `camera-profiles.json` speaker entries | Correct speaker names assigned to segments matching current pipeline output on fixtures | +| F3.4 | Transcript editing logic | Port of cut derivation and sentence merging from `scripts/edit-transcript.js`; `PAUSE_THRESHOLD`, `WORD_DURATION_ESTIMATE`, `CUT_START_BIAS` constants preserved exactly | Output transcript matches current `edit-transcript.js` output on real fixtures; constants not altered | +| F3.5 | Face detection subprocess | `tokio::process` call to mediapipe; output parsed into per-speaker closeup viewport | Viewport output matches current `setup-camera.js` output on fixtures | +| F3.6 | Thumbnail background removal subprocess | `tokio::process` call to rembg | Output matches current rembg output on fixtures | +| F3.7 | Short-form clip extraction | Short-form clips extracted as a first-class pipeline output from the same long-form source — not a separate wizard pass; produces `transcript.json` with `meta.outputAspect: "9:16"` and correct `meta.videoStart`/`meta.videoEnd` | 3–5 clips produced in the same pipeline run as the long-form edit; no second transcription or sync pass required | + +--- + +## Epic 4: Multi-brand Support + +**Opportunity addressed:** [Opportunity 06 — Multi-brand support](06-opportunities/06-multi-brand-support.md) + +**RFC reference:** Not an explicit build-order step, but brand config must be in the schema from the start — retrofitting it later (as the current codebase is doing in Phase 0.5) is exactly what the rewrite should avoid. **Build alongside Epic 1** so the compositor, schema types, and asset loading are brand-aware from the first render. + +**Problem it solves:** ragTech brand assets are hardcoded throughout the current pipeline. Delivering client footage with a different brand is not possible without manual file replacement. The service business (studio-rental partnership) cannot operate without brand-per-job support. + +**Definition of done:** A pipeline run for Brand A and Brand B on the same raw footage produces two fully different branded outputs — logos, colors, fonts, host name cards, intro/outro, audio — with no manual file-swapping. Adding a third brand requires only a new brand directory and config file, zero code changes. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F4.1 | Brand config schema | Versioned JSON schema for per-job brand input: logo, colors, fonts, hosts (with camera angle mapping), intro/outro music, background music, mascot/visual assets, overlay template selection | Schema round-trips without loss; ragTech brand expressed as Brand 0 using this schema, not as a special case | +| F4.2 | Brand asset loader | Resolves brand directory at runtime from config; loads all assets for the job | Brand A and Brand B assets load independently in the same process; wrong-brand asset cannot be loaded for a given job | +| F4.3 | Core overlay template library | Parameterised intro sequence, name card/lower-third, and outro — driven by brand config (colors, fonts, logo, host images) | Same template renders correctly with two different brand configs; no brand-specific code paths inside the template | +| F4.4 | Brand registry | Convention: adding a brand = adding `brands/{name}/brand.json` + `brands/{name}/assets/`; no code changes required | Third brand added by directory + config only; pipeline picks it up without a code change or rebuild | +| F4.5 | ragTech brand migrated as Brand 0 | ragTech brand expressed entirely via the brand config schema — no ragTech-specific code remaining in the compositor or overlay logic | ragTech episode renders identically before and after migration; Techybara mascot, Nunito font, intro/outro music all driven from `brands/ragtech/brand.json` | + +--- + +## Epic 5: Visual Transcript Editor + +**Opportunities addressed:** [Opportunity 05 — Visual transcript editor](06-opportunities/05-visual-transcript-editor.md), [Opportunity 04 — Transcript token timing accuracy](06-opportunities/04-transcript-timing-accuracy.md) (visual cut placement component) + +**RFC reference:** Build Order Step 4. Built last — after the compositor API (Epic 1) is stable — so the preview surface shares the exact same code path as final render. + +**Problem it solves:** The current editing interface (plain text file in VSCode with custom cue syntax) requires editors to make timing-sensitive cut decisions blind — no video preview, no waveform, no playback. Hook clip boundaries require typing estimated float timestamps. Cue syntax is not discoverable. The gap between editing and seeing the result drives most of the trial-and-error in the current workflow. + +**Definition of done:** An editor can complete a full episode edit — including hook clip in/out points — without typing a single float timestamp, and can preview any cut in-context before committing to a full render. In-editor preview matches final render output frame-for-frame. + +### Features + +| ID | Feature | Description | Acceptance criteria | +|----|---------|-------------|---------------------| +| F5.1 | egui timeline with waveform and playhead | Immediate-mode timeline rendering: waveform display, playhead scrubbing, clip lane visualisation | Waveform renders at 60fps without frame drops; playhead scrubs smoothly across a 60-minute episode | +| F5.2 | Transcript ↔ timeline sync | Clicking a word in the transcript scrubs the video to that word's `t_dtw` timestamp; moving the playhead highlights the corresponding word | Both directions work with <1 frame latency; clicking a word in a cut segment shows the cut state visually | +| F5.3 | Visual cut placement | Dragging a cut boundary on the timeline updates `t_dtw`/`t_end` in the transcript; the transcript text reflects the cut without manual float entry | Cut boundary placed visually lands within 1 frame of intended point; corresponding token in transcript updated correctly; no float typing required | +| F5.4 | Hook clip in/out point setting | In/out points for hook clips (`hookFrom`/`hookTo`) set by marking a range on the timeline | Hook boundary set visually; `hookFrom`/`hookTo` in transcript updated correctly; no float estimation | +| F5.5 | In-context preview | Selected cut or clip range plays back immediately in the editor using the compositor — same code path as final render | Preview output matches final render output frame-for-frame for the same segment; no separate render step to verify a cut | +| F5.6 | Camera angle preview per segment | Shows the active camera angle and viewport crop for a given segment inline in the timeline | Correct angle and crop displayed for each segment; switching angles visible without leaving the editor | + +--- + +## Sequencing summary + +``` +[Spike] FFmpeg vs gstreamer-rs ← runs in deckcreate repo; parallel with Epic 0 +[Epic 0] Project Setup & AI Harness ← new repo; fixtures + baselines + harness; must finish before Epic 1 + ↓ +[Epic 1] Native Render Engine ← compositor API must be stable before anything builds on it +[Epic 4] Multi-brand Support ← parallel with Epic 1; brand schema must be in from day 1 + ↓ +[Epic 2] Transcription and Alignment ← begin once Epic 1 compositor API is stable +[Epic 3] Pipeline Orchestration ← parallel with Epic 2; wires around the proven render core + ↓ +[Epic 5] Visual Transcript Editor ← built last; shares compositor from Epic 1 +``` + +The spike and Epic 0 are the only things that can run in parallel at the start — everything else has a hard dependency chain. Epics 1+4 are parallel. Epics 2+3 are parallel after Epic 1. Epic 5 is last. From 4dac352c1d4c942942a511cbce35a043161e98b1 Mon Sep 17 00:00:00 2001 From: Natasha Ann Date: Mon, 20 Jul 2026 11:29:38 +0800 Subject: [PATCH 7/7] docs(epics): update epics with spike findings --- docs/product/07-epics.md | 49 +++++++++++++++++++++------------------- 1 file changed, 26 insertions(+), 23 deletions(-) diff --git a/docs/product/07-epics.md b/docs/product/07-epics.md index 728f2f0..f3361c3 100644 --- a/docs/product/07-epics.md +++ b/docs/product/07-epics.md @@ -6,27 +6,27 @@ This document is the direct bridge from product discovery to engineering. Each e - Each epic becomes a GitHub Epic (or milestone, depending on how the new repo structures work). - Each feature under an epic becomes one or more GitHub issues, with acceptance criteria derived from the success metrics in the linked opportunity doc. - Epic 0 must be the first epic completed — its outputs (test fixtures, golden-frame baselines, AI harness) are the acceptance criteria infrastructure that every subsequent epic depends on. -- The spike must be resolved before Epic 1 issues are written — its outcome determines which FFmpeg binding the entire render engine is built against. It can run in parallel with Epic 0 since it lives in the current repo. +- The spike is resolved — `ffmpeg-the-third` is the chosen binding (see `spike/rust-decode-spike/FINDINGS.md`). Epic 1 issues can be written now, with F1.0 (LGPL FFmpeg build) as the first task. --- ## Pre-epic: FFmpeg vs gstreamer-rs spike -**Not an epic — a one-off technical investigation that must complete before Epic 1 begins.** +**Status: RESOLVED.** See [`spike/rust-decode-spike/FINDINGS.md`](../../spike/rust-decode-spike/FINDINGS.md) — verified on Mac M3, Mac M2, and Windows/NVIDIA. Decision feeds into Epic 1 F1.0. -The RFC calls this out explicitly as non-blocking to the plan but blocking to the render engine epic: the decode/encode binding choice determines every subsequent API call in the compositor. Deciding mid-epic would force a rewrite of already-written code. +**Decision: `ffmpeg-the-third`** (the actively-maintained fork of `ffmpeg-next`; `ffmpeg-next` itself is maintenance-mode only — confirmed by the spike). -**Question to answer:** Is `ffmpeg-next` (direct FFmpeg C API bindings) or `gstreamer-rs` the better foundation for the render engine on the target hardware matrix (Mac M2/M3 VideoToolbox, Windows NVIDIA NVENC/NVDEC)? +**Rationale summary:** +- On Mac M2 and M3: `ffmpeg-the-third` with VideoToolbox hardware path is pixel-exact (mean diff = 0.000, max = 0) against reference frames at both tested timestamps. `gstreamer-rs`/`vtdec_hw` fails pixel tolerance on both chips (mean ~3.86, max ~168–181) — root cause is a colorimetry/YUV-matrix mismatch during GPU→CPU readback, not a seek bug. +- On Windows/NVIDIA: both candidates' hardware paths are active but fail the spike's pixel tolerance. `ffmpeg-the-third`/CUDA is materially closer to reference than `gstreamer-rs`/`nvh264dec` (lower mean and much lower max diff). Both require follow-up hardening in Epic 1. +- `ffmpeg-the-third`'s missing high-level hwaccel API requires ~40 lines of unsafe FFI for the VideoToolbox and CUDA paths — this is a one-time bounded cost, and the spike's `hw` module (`ffmpeg-candidate/src/main.rs`) is the working reference implementation. Do not re-derive from FFmpeg's C `hw_decode.c` example — use the spike code directly. +- `gstreamer-rs` has the easier integration surface but unresolved hardware-path pixel failures on both Mac and Windows, plus a Cocoa run-loop dependency (`NSApplication`) introduced by `vtdec_hw`'s GL-memory output — a real open question for a headless render engine. -**Evaluation criteria:** -- Hardware codec access: can it reach VideoToolbox on Mac and NVENC/NVDEC on Windows without extra shims? -- Seek/decode accuracy: frame-accurate seek to an arbitrary timestamp (required for jump-cut editing). -- Build complexity: does it cross-compile cleanly on all three target platforms from CI? -- Maintenance: crate maturity, update frequency, known soundness issues. +**Open questions carried into Epic 1** (not re-investigated in the spike): +1. Windows/NVIDIA hardware decode pixel quality needs root-cause and correction — treat as integration-incomplete until resolved in Epic 1. +2. The GPL/LGPL FFmpeg build configuration is a **hard prerequisite for Epic 1**: both Mac and Windows test machines used Homebrew/system FFmpeg builds with `--enable-gpl --enable-libx264 --enable-libx265`, making the software `avdec_h264` decode path GPL-contaminated. A clean LGPL-only FFmpeg build (without `--enable-gpl`) fully supports H.264 decode and removes this risk — this must be resolved before Epic 1 begins, not deferred. See F1.0 below. -**Spike output:** a short findings doc (not a polished RFC — a decision record with the test results) committed to the spike branch. The finding becomes the `encoderProfile` design decision that Epic 1 is built against. - -**Where to run it:** a `spikes/rust-decode-spike/` directory in the current repo (deckcreate), on its own branch. Findings port to the new repo; the spike code does not. +**Where the spike ran:** `spikes/rust-decode-spike/` in the deckcreate repo. Spike code is not ported — only the `hw` module pattern and the findings doc carry forward. --- @@ -48,8 +48,8 @@ The RFC calls this out explicitly as non-blocking to the plan but blocking to th | F0.2 | CLAUDE.md — AI harness | Agent implementation convention for the new repo: crate layout, how to run tests (`cargo test`, `cargo clippy`, `cargo fmt --check`), implementation doc format (one doc per multi-step task in `docs/implementation-guides/`), status check per step, one commit per step, per-PR checklist adapted for Rust | An agent given a feature issue can orient, find the right crate, run the test suite, and know what done looks like without any out-of-repo context | | F0.3 | CI/CD cross-platform matrix | GitHub Actions workflow: build + test on Mac ARM (M-series), Windows x64 with NVIDIA GPU runner; clippy and rustfmt checks on every PR; test matrix fails fast if any target fails | Green CI required to merge any PR; a failure on Windows/NVIDIA is as blocking as a failure on Mac | | F0.4 | Test fixture export | Export real `transcript.json` + `camera-profiles.json` from at least two already-shipped episodes from the current production pipeline; commit as versioned fixtures in `tests/fixtures/` | Fixtures round-trip through the new repo's Rust schema structs without loss; used as the canonical input for all Epic 1–3 acceptance tests | -| F0.5 | Golden-frame baseline export | For each fixture episode, render a reference set of frames using the current Remotion pipeline and commit as baselines in `tests/baselines/`; include a PSNR threshold and audio hash per fixture | Epic 1's compositor output can be diffed against these baselines with a single test command; baselines are reproducible from the current pipeline if they need to be regenerated | -| F0.6 | Shared test utilities | Common test helpers in `crates/test-utils`: fixture loader, PSNR calculator, audio hash utility, frame extractor — consumed by all epic test suites | Any crate's test suite can load a fixture and run a PSNR diff in under 5 lines of test code | +| F0.5 | Compositor property specs and sanity frames | **Tier 1 (primary gate):** a `properties.json` spec per fixture episode, generated from the deckcreate pipeline's JS logic, defining expected compositor state (active angle, viewport crop, cut/hook/speaker flags, source timestamp) at key checkpoints. **Tier 2 (sanity signal only):** a small set of Remotion-rendered frames per fixture for gross visual regression detection — not a pass/fail gate. | Tier 1 property specs are the acceptance criterion for Epic 1 correctness; Tier 2 frames are informational only. Remotion pixel output is not the correctness target — deckcreate artifacts were produced as prototypes, not pixel-perfect benchmarks. | +| F0.6 | Shared test utilities | Common test helpers in `crates/test-utils`: fixture loader, property spec checker (loads `properties.json` and asserts compositor state), frame extractor, audio hash utility — consumed by all epic test suites | Any crate's test suite can load a fixture and assert compositor state against its property spec in under 5 lines of test code | | F0.7 | Dev tooling | `rustfmt.toml`, `.clippy.toml`, pre-commit hooks (or equivalent) enforcing format and lint on changed files; `Makefile` or `justfile` with canonical commands (`make test`, `make lint`, `make baseline`) | A new contributor (or agent) can set up the dev environment and run the full check suite with documented commands; no manual tool invocation required | --- @@ -62,19 +62,20 @@ The RFC calls this out explicitly as non-blocking to the plan but blocking to th **Problem it solves:** The current Remotion/headless-Chromium render path has a hard 2–5fps ceiling with no GPU path, producing 6–18 hour renders for a 60-minute episode. Hardware inconsistency (Windows/NVIDIA silently falling back to software encode) means output quality depends on which machine ran the job. -**Definition of done:** A render of the same episode fixture on Mac M2, Mac M3, and Windows/NVIDIA produces output that passes a PSNR/hash diff gate and completes in under 1 hour — verified across all three hardware targets. +**Definition of done:** A render of the same episode fixture on Mac M2, Mac M3, and Windows/NVIDIA produces output that passes the Tier 1 compositor property spec checks (from F0.5) and completes in under 1 hour — verified across all three hardware targets. Windows NVDEC pixel drift (open from spike) must be root-caused and corrected before this epic closes. ### Features | ID | Feature | Description | Acceptance criteria | |----|---------|-------------|---------------------| -| F1.0 | Spike outcome applied | FFmpeg binding (or gstreamer-rs) chosen from the spike; dependency declared in `Cargo.toml` | Spike findings doc exists and is linked from this epic's first issue | +| F1.0 | GPL/LGPL FFmpeg build prerequisite | Before any Epic 1 code is written: verify a clean LGPL-only FFmpeg build (compiled without `--enable-gpl --enable-libx264 --enable-libx265`) is available on all CI and dev machines. This removes GPL contamination from the `avdec_h264` software decode path. The spike used Homebrew/system FFmpeg with GPL flags — this must be resolved before the `ffmpeg-the-third` binding is linked in production builds. Document the required build flags in `docs/implementation-guides/ffmpeg-build.md`. | CI pipeline uses an LGPL-clean FFmpeg on Mac ARM and Windows x64; `ffmpeg-the-third` links against it without GPL symbols; build flags documented and reproducible. | +| F1.0a | `ffmpeg-the-third` binding bootstrapped | Add `ffmpeg-the-third` to `crates/compositor/Cargo.toml`; port the spike's `hw` module (`spike/rust-decode-spike/ffmpeg-candidate/src/main.rs`) as the starting point for the hardware decode path — do not re-derive from FFmpeg's C `hw_decode.c` example. VideoToolbox and CUDA paths brought over from the spike. | `cargo build --workspace` succeeds on Mac and Windows with `ffmpeg-the-third` declared; hw module compiles against the LGPL-clean build; spike findings doc linked from this feature's issue. | | F1.1 | Multi-angle compositor | wgpu-based compositor that stacks multiple decoded video planes, applies viewport crops, and composites to an output surface — same code path for preview and final render | Composites 3 angles with correct viewport transforms matching the `camera-profiles.json` spec; verified against golden frames from the current Remotion output | -| F1.2 | VideoToolbox decode/encode path | Hardware-accelerated H.264 decode and encode via AVFoundation + VideoToolbox on macOS | Render on Mac M2 and Mac M3 completes in under 1 hour for a 60-min episode; output passes PSNR gate vs Remotion baseline | -| F1.3 | NVENC/NVDEC decode/encode path | Hardware-accelerated H.264 decode (NVDEC) and encode (NVENC) on Windows/NVIDIA | Render on Windows/NVIDIA completes in under 1 hour; output passes PSNR gate; no silent fallback to libx264 | +| F1.2 | VideoToolbox decode/encode path | Hardware-accelerated H.264 decode and encode via AVFoundation + VideoToolbox on macOS. **Spike result:** `ffmpeg-the-third` VideoToolbox path is pixel-exact on both Mac M2 and Mac M3 (mean diff = 0.000, max = 0 at both tested timestamps) — the implementation baseline exists in the spike's `hw` module. | Render on Mac M2 and Mac M3 completes in under 1 hour for a 60-min episode; output passes Tier 1 property spec checks; VideoToolbox path engaged (verified in logs, not silent software fallback). | +| F1.3 | NVENC/NVDEC decode/encode path | Hardware-accelerated H.264 decode (NVDEC) and encode (NVENC) on Windows/NVIDIA. **Spike result:** NVDEC hardware path is active on both candidates but both fail the spike's pixel tolerance (`ffmpeg-the-third`/CUDA mean diff ≈ 1.85 / max 91; gstreamer-rs/nvh264dec mean ≈ 3.86 / max 168). Root cause is unknown — this is an open Epic 1 investigation item, not a deferred nice-to-have. Both the LGPL build (F1.0) and the root-cause investigation must complete before this feature can be marked done. | Root cause of Windows NVDEC pixel drift identified and corrected; render on Windows/NVIDIA completes in under 1 hour; output passes Tier 1 property spec checks; no silent fallback to libx264; NVDEC hardware engagement verified in logs. | | F1.4 | Real hardware capability probing | Replaces the `process.platform`/`process.arch` string heuristic in `hardware.ts` with actual GPU/codec capability queries at startup | Correctly identifies VideoToolbox on Mac and NVENC on Windows/NVIDIA; logs the resolved encoder profile at startup; does not silently fall back | | F1.5 | Unified encoder profile mapping | Single `encoderProfile → FFmpeg/codec args` mapping consumed by every pipeline stage — no per-script encoder decisions | No pipeline script sets its own encoder; all route through the profile mapping; adding a new profile requires one code change in one file | -| F1.6 | Golden-frame validation suite | Automated diff of new-engine output against Remotion output for real-episode fixtures — PSNR/hash diff for video, audio-hash diff for audio | Suite runs on CI across Mac M2, Mac M3, and Windows/NVIDIA; fails if PSNR drops below threshold or audio hash mismatches | +| F1.6 | Property spec validation suite | Automated Tier 1 property spec checks (active angle, viewport crop, cut/hook/speaker flags, source timestamps) against `properties.json` fixtures for all fixture episodes; Tier 2 sanity frames generated as informational artifact. | Suite runs on CI across Mac M2, Mac M3, and Windows/NVIDIA; Tier 1 checks are the pass/fail gate; Tier 2 frames stored as CI artifacts for visual review — not a blocking gate. | | F1.7 | `transcript.json` + `camera-profiles.json` schema ingestion | Rust structs for both schemas; round-trips real fixture files without loss | Schema conformance tests pass on all fixture files exported from the current production pipeline | --- @@ -171,10 +172,10 @@ The RFC calls this out explicitly as non-blocking to the plan but blocking to th ## Sequencing summary ``` -[Spike] FFmpeg vs gstreamer-rs ← runs in deckcreate repo; parallel with Epic 0 -[Epic 0] Project Setup & AI Harness ← new repo; fixtures + baselines + harness; must finish before Epic 1 +[Spike] FFmpeg vs gstreamer-rs ← RESOLVED: ffmpeg-the-third chosen; findings in spike/rust-decode-spike/FINDINGS.md +[Epic 0] Project Setup & AI Harness ← new repo; fixtures + property specs + harness; must finish before Epic 1 ↓ -[Epic 1] Native Render Engine ← compositor API must be stable before anything builds on it +[Epic 1] Native Render Engine ← starts with F1.0 (LGPL FFmpeg build) then F1.0a (hw binding bootstrap) [Epic 4] Multi-brand Support ← parallel with Epic 1; brand schema must be in from day 1 ↓ [Epic 2] Transcription and Alignment ← begin once Epic 1 compositor API is stable @@ -183,4 +184,6 @@ The RFC calls this out explicitly as non-blocking to the plan but blocking to th [Epic 5] Visual Transcript Editor ← built last; shares compositor from Epic 1 ``` -The spike and Epic 0 are the only things that can run in parallel at the start — everything else has a hard dependency chain. Epics 1+4 are parallel. Epics 2+3 are parallel after Epic 1. Epic 5 is last. +The spike is resolved. Epic 0 is the current blocker — until fixtures (F0.4), property specs (F0.5), and the AI harness (F0.2) are complete, Epic 1 acceptance criteria cannot be evaluated. Epics 1+4 are parallel once Epic 0 is done. Epics 2+3 are parallel after Epic 1. Epic 5 is last. + +**Epic 1 internal sequencing note:** F1.0 (LGPL FFmpeg build verification) must be the first task in Epic 1 — it is a prerequisite for F1.0a (binding bootstrap) and all subsequent F1.x work. F1.3 (Windows NVDEC) has an open root-cause investigation; it is not parallel-safe with F1.2 until the spike pixel drift is understood.