Science Explained/Brief
Synthetic office-email benchmark leaves AI agents short of 80%
A preprint describes an automated pipeline that generates synthetic workplace-email datasets and reports that standard agentic baselines score below 80% on them. The evidence is a preprint built on simulated data, so the numbers describe one constructed test rather than real enterprise performance.
BriefPublished 14 September 20261 min read1 linked source · 3 checked facts
A preprint describes an automated pipeline, WinSyn, that builds synthetic datasets of workplace emails, together with long- and short-form questions and answers grounded in those emails. The simulated projects run for several months and involve up to 25 employees across multiple roles, with the data designed to carry ambiguity and information spread across messages.
The authors evaluated standard agentic baselines on the datasets using recent frontier models. Aggregate scores averaged over all queries stayed below 80% on each dataset. The paper's own conclusion is that this suggests more work remains before enterprise deployment.
Our view
This is a benchmark-building exercise on synthetic data, so the sub-80% scores describe how today's agents handle one constructed test, not how they perform in real enterprises.
What the reporting says: Aggregate scores averaged over all queries remain below 80% for each dataset, and that these findings "suggest that more work remains to be done for enterprise deployment."