Science Explained/Brief
Preprint tests LLM judge for comparing clinical timelines to case reports
A preprint on arXiv describes GAVEL, a protocol that uses a large language model to compare two clinical timelines against the original case report. In tests across 126 reports, merged timelines were preferred in 77.0% of comparisons and reduced discrepancies from 7.63 to 0.85 per report.
BriefPublished 15 September 20261 min read1 linked source · 5 checked facts
The work is a preprint, meaning it has not yet been peer-reviewed. The authors evaluated an event matcher, reviewed 2,738 findings from two models, and ranked six LLM extractors and two human annotators. Manual review confirmed 89.4% and 88.6% of findings. The authors say GAVEL supports report-based comparison and revision without treating either timeline as ground truth.
Our view
This preprint offers a method for evaluating timeline extraction without a perfect reference, but its results are preliminary and have not been peer-reviewed.
What the reporting says: GAVEL is an LLM judge protocol that compares two timelines with the case report and returns a discrepancy type, verdict, and report passage for each difference.