Science Explained/Brief
CADWorld benchmark tests AI agents on long-horizon FreeCAD design tasks
A benchmark called CADWorld evaluates computer-use agents on long-horizon mechanical CAD tasks in FreeCAD, with success measured by executable checks on saved artifacts. The strongest of seven agents solved 17.5% of tasks, against an 87.0% expert reference pass.
BriefPublished 16 September 20261 min read1 linked source · 4 checked facts
A new benchmark called CADWorld tests computer-use agents on long-horizon mechanical design in FreeCAD. It includes 200 tasks across 11 workflow categories, from sketching and part modeling to assembly, CAM, FEM, measurement, mesh processing, and technical drawing. Agents work only through screenshots and GUI actions, and success is judged by executable checks on saved FreeCAD artifacts and auxiliary outputs.
Across seven current agents, the strongest reached 17.5% success, while an expert reference pass was 87.0%. The authors report that weaker agents often fail before producing a valid artifact, whereas stronger agents more often fail on structural, geometric, and construction-process requirements.
Our view
The benchmark exposes a gap between general GUI competence and reliable execution of persistent engineering workflows, but these are early results from a single benchmark and not a settled finding.
What the reporting says: Across seven current agents on the full benchmark, the strongest agent achieves 17.5% success, compared with an 87.0% expert reference pass.