A preprint describes an automated pipeline, WinSyn, that builds synthetic datasets of workplace emails, together with long- and short-form questions and answers grounded in those emails. The simulated projects run for several months and involve up to 25 employees across multiple roles, with the data designed to carry ambiguity and information spread across messages.

The authors evaluated standard agentic baselines on the datasets using recent frontier models. Aggregate scores averaged over all queries stayed below 80% on each dataset. The paper's own conclusion is that this suggests more work remains before enterprise deployment.