The abstract reports that GLARE achieved average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN but remaining below human continuations. The benchmark includes 24,794 future-facing queries. As a preprint, these results have not been independently verified.
Science Explained/Brief
Vendor-Native Coding Harnesses Show No Clear Average Advantage in Paired Test
A new arXiv paper compares agentic coding harnesses paired with the same models on a private, contamination-controlled suite. It reports no resolved average advantage…
