The setup: an AI agent drafts a patent, a separately invoked AI judge scores the draft and returns structured feedback, and the agent revises. Across several inventions and drafting-agent configurations, this loop consistently improved the judge's own quality scores, while revision without judge feedback tended to plateau. The authors also report that iterative judge feedback let a lower-reasoning agent approach the performance of a substantially more expensive high-reasoning one.

The check: the judge was compared with an independent evaluation by a professional patent attorney. Agreement was meaningful but strongly metric-dependent, with systematic calibration differences. That is the boundary to hold on to — the improvement is largely measured by the judge being tested, not by a legal standard.