Observed during: Tragedy Practice Mode planning (2026-07-17).
Why it matters: A fresh blind benchmark: one song segment, 3 distinct reference dancers, consented practice attempts, expert-marked critical moments. Measures whether NodeVideo keeps interpretations distinct and whether coaching improves the next attempt (the metric that matters). Also the first real data for the fail-closed calibration gate, which has never had a manifest.
Why out of scope: Needs consented multi-dancer footage and expert labels; P0 is single-reference.
Risk: B.
Observed during: Tragedy Practice Mode planning (2026-07-17).
Why it matters: A fresh blind benchmark: one song segment, 3 distinct reference dancers, consented practice attempts, expert-marked critical moments. Measures whether NodeVideo keeps interpretations distinct and whether coaching improves the next attempt (the metric that matters). Also the first real data for the fail-closed calibration gate, which has never had a manifest.
Why out of scope: Needs consented multi-dancer footage and expert labels; P0 is single-reference.
Risk: B.