The antecedent annotator (#829) will produce a gold corpus of labeled Id./supra → antecedent links. Before v1.0 ships, we need a deliberate decision on how that corpus is released, since publication is not reversible.
Options, in no particular order:
- Full release — publish all labels (e.g. CC-BY) alongside the repo's existing committed corpus fixtures
- Benchmark split — publish a versioned dev/test benchmark subset; keep the remainder as a held-out evaluation set
- Defer — keep the corpus unpublished for now and revisit after v1.0
Considerations to settle as part of the decision: license for the labels (source opinions are public-domain via CourtListener), dev/test split discipline if any subset is published, whether the resolver-accuracy CI net needs public data to run on fork PRs, and versioning of the corpus artifact.
No action needed until v1.0 launch planning; filing now so the decision point doesn't get lost.
🤖 Generated with Claude Code
The antecedent annotator (#829) will produce a gold corpus of labeled Id./supra → antecedent links. Before v1.0 ships, we need a deliberate decision on how that corpus is released, since publication is not reversible.
Options, in no particular order:
Considerations to settle as part of the decision: license for the labels (source opinions are public-domain via CourtListener), dev/test split discipline if any subset is published, whether the resolver-accuracy CI net needs public data to run on fork PRs, and versioning of the corpus artifact.
No action needed until v1.0 launch planning; filing now so the decision point doesn't get lost.
🤖 Generated with Claude Code