Gate A's headline (+0.065 LODO increment over logprob-only) uses the committed logprob feature set. A follow-up analysis on related data found that a stronger output-confidence baseline (min token logprob over the answer) closes much of the gap in some settings.
Task: rerun the Gate A LODO comparison with a min-token-logprob feature added to the baseline set, on the published Stage 1 features. If the increment shrinks, that is a result we want on record; if it survives, the claim gets stronger. CPU-only.
Do not touch campaign/frozen/ — this is a new analysis, not a re-fit of frozen artifacts.
Gate A's headline (+0.065 LODO increment over logprob-only) uses the committed logprob feature set. A follow-up analysis on related data found that a stronger output-confidence baseline (min token logprob over the answer) closes much of the gap in some settings.
Task: rerun the Gate A LODO comparison with a min-token-logprob feature added to the baseline set, on the published Stage 1 features. If the increment shrinks, that is a result we want on record; if it survives, the claim gets stronger. CPU-only.
Do not touch campaign/frozen/ — this is a new analysis, not a re-fit of frozen artifacts.