Everything published here is AUROC (ranking). For routing decisions you also want calibration: reliability curves and ECE for the frozen classifiers, per source, in-domain vs zero-shot transfer.
CPU-only: score stage2_features.jsonl with campaign/frozen/ (see campaign/reproduce_mini.py for the loading pattern) and plot reliability diagrams. Well-calibrated in-domain but miscalibrated on transfer would itself be a finding worth writing up.
Everything published here is AUROC (ranking). For routing decisions you also want calibration: reliability curves and ECE for the frozen classifiers, per source, in-domain vs zero-shot transfer.
CPU-only: score stage2_features.jsonl with campaign/frozen/ (see campaign/reproduce_mini.py for the loading pattern) and plot reliability diagrams. Well-calibrated in-domain but miscalibrated on transfer would itself be a finding worth writing up.