The pitch is about episode economics rather than peak capability. Co-work agents chain information gathering, tool use, coding and file manipulation across many model invocations, so cost and latency pile up over the whole run. The authors argue many everyday steps are about state tracking, coordination, recovery and follow-through, not frontier-scale reasoning.
The one useful detail for operators: the paper says its aggregate score across four representative benchmarks puts it at the low-cost knee of the observed cost-performance Pareto frontier, under its own stated evaluation and pricing protocol. Supporting evaluations in tool calling, coding and instruction following are reported to preserve broad agentic capability.
Evidence is thin on the operational specifics. The abstract does not give per-benchmark numbers, latency figures, harness configurations or deployment requirements, so treat the frontier claim as the authors' framing rather than a measured result you can plan against.
