The abstract reports that GLARE achieved average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN but remaining below human continuations. The benchmark includes 24,794 future-facing queries. As a preprint, these results have not been independently verified.