Existing benchmarks for enterprise question-answering often have limited real-world complexity, short-form responses, and unnatural queries, according to the paper. The new pipeline generates synthetic emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. It simulates projects spanning several months and involving up to 25 interacting employees across multiple roles, emphasizing ambiguity and distributed information. When evaluated with standard agentic baselines using frontier models, aggregate scores remained below 80% for each dataset.
Practical AI/Report
New method explains self-driving car AI decisions in real time
Researchers developed CW-Net, a method that translates an autonomous vehicle AI planner’s reasoning into understandable concepts and outputs explanations alongside its trajectory. Road tests…
