Balanced accuracy ranged from 0.156, below chance, to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest reasoning effort level, while Gemini 3 Pro stayed above 0.94 at every level. Gemini 3 Pro at low reasoning effort outscored GPT-5.2 at xhigh reasoning effort for roughly 5% of the cost per trial. When the meta-evaluator erred, it tended to rely on surface cues rather than operational reasoning. The abstract advises context engineers to make the intended referent explicit at each stage.
Science Explained/Brief
Vendor-Native Coding Harnesses Show No Clear Average Advantage in Paired Test
A new arXiv paper compares agentic coding harnesses paired with the same models on a private, contamination-controlled suite. It reports no resolved average advantage…
