Science Explained/Brief
AI judges improve AI-written patent drafts, but match a human attorney only partly
A new arXiv preprint describes a patent-drafting testbed in which a separate AI judge scores drafts and feeds back revisions. The reported gains are measured largely by that same judge, and its agreement with a professional patent attorney was meaningful but depended heavily on which metric was used.
BriefPublished 15 September 20261 min read1 linked source · 4 checked facts
The setup: an AI agent drafts a patent, a separately invoked AI judge scores the draft and returns structured feedback, and the agent revises. Across several inventions and drafting-agent configurations, this loop consistently improved the judge's own quality scores, while revision without judge feedback tended to plateau. The authors also report that iterative judge feedback let a lower-reasoning agent approach the performance of a substantially more expensive high-reasoning one.
The check: the judge was compared with an independent evaluation by a professional patent attorney. Agreement was meaningful but strongly metric-dependent, with systematic calibration differences. That is the boundary to hold on to — the improvement is largely measured by the judge being tested, not by a legal standard.
Our view
The gains are real inside the paper's own scoring loop, but the attorney comparison suggests the judge is not yet a substitute for professional judgment.
What the reporting says: Judge-guided revision consistently improves judge-assessed quality while unguided revision tends to saturate, and that validation against a professional patent attorney found meaningful but strongly metric-dependent agreement and systematic calibration differences.