Skip to content

Calibration analysis: are the scores calibrated, or just ranked? #12

Description

@solarkyle

Everything published here is AUROC (ranking). For routing decisions you also want calibration: reliability curves and ECE for the frozen classifiers, per source, in-domain vs zero-shot transfer.

CPU-only: score stage2_features.jsonl with campaign/frozen/ (see campaign/reproduce_mini.py for the loading pattern) and plot reliability diagrams. Well-calibrated in-domain but miscalibrated on transfer would itself be a finding worth writing up.

Metadata

Metadata

Assignees

No one assigned

    Labels

    analysis ideaNew analysis over the published traces/features — usually no GPU neededgood first issueGood for newcomers

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions