CreatorBench v1.4: adaptive routing, sealed evidence, and review UI#33
Draft
HomenShum wants to merge 25 commits into
Draft
CreatorBench v1.4: adaptive routing, sealed evidence, and review UI#33HomenShum wants to merge 25 commits into
HomenShum wants to merge 25 commits into
Conversation
Owner
Author
|
Final preflight preview (deployment-protected):
Signed-in browser verification: real report rendered, no console errors, UTF-8, no horizontal overflow, null correction/cost metrics fail closed as “Not reported.” All CI gates are green. The hard benchmark-size gate intentionally remains red (5/250 clips, 5/75 creators, 4/15 domains, 40/2,000 instances); no blinded human-reviewed renders exist, so usable output remains 0 and no arbitrary-footage performance claim is made. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
CreatorBench v1.4 is the first frozen NodeVideo benchmark release to meet the declared corpus scale and workflow-coverage targets while preserving creator/source-disjoint splits, sealed private evaluation, public proof isolation, and honest claim generation.
This PR deliberately does not claim universal editing performance. It measures what the current orchestration layer can defend: rights and policy checks, route eligibility, review/abstention behavior, canonical workflow execution evidence, deterministic export/reopen integrity, and the remaining need for blinded human editing-quality review.
What changed
Frozen v1.4 corpus
Speech metadata covers 92/134 eligible sources with 3,422 transcribed words, 717 quote segments, and 286 detected silence regions.
Freeze and sealed evidence
creatorbench-v1.4creatorbench-freeze:f3fb72fb06287374e5206fee95850b605cab74946a5cb422fef94ff6858fe5fdsha256:f2b9b95d371ab1b27605842dc473b59b1df92116fd9f050ad80a73dbf3b53d3esha256:f9c4a4e4b5e3b2447914066686125a0892c2a73da18b44edc2a31f4aec07f00fSealed private routing
Canonical public claim:
Deterministic workflow evidence
Talking-head public pilot:
Public multi-format center-crop pilot:
These are technical workflow checks, not human editing-quality claims.
Verification
npm run check: passProduction proof
804ec170d9ad44c928a628e5c46518bc38e7682espotted-cat-868d3d69511e45ba9197b230616d9d820f774d8caab7a9b96fdbd6b64793735ae0aThe frozen source commit fixes the benchmark population and evaluator boundary. Later commits publish generated evidence, bind the UI tests to v1.4, and stamp the deployment without changing the frozen private population.
Remaining hard gate
The benchmark infrastructure, corpus, workflow coverage, sealed execution, UI, and production proof are complete. The release still needs real blinded human judgments and correction-time measurements before it can report:
No synthetic or model-generated labels are substituted for that human work.
Claim boundary
Defensible now:
Not defensible yet:
That stronger statement requires completing the blinded human-review phase and passing prespecified workflow and subgroup thresholds.