In the Gate D table, bfcl shows a negative increment (-0.020). Tool-calling failure is a different beast from factual error — the interesting question is whether the negative increment is uniform or concentrated (specific tool types? argument-fabrication vs wrong-tool-choice? fixed-truth-style subsets?).
CPU-only over the published Stage 2 features and traces. This connects to the broader question the repo cares about: what is the signal actually reading when it stops predicting error.
In the Gate D table, bfcl shows a negative increment (-0.020). Tool-calling failure is a different beast from factual error — the interesting question is whether the negative increment is uniform or concentrated (specific tool types? argument-fabrication vs wrong-tool-choice? fixed-truth-style subsets?).
CPU-only over the published Stage 2 features and traces. This connects to the broader question the repo cares about: what is the signal actually reading when it stops predicting error.