The work is a preprint, meaning it has not yet been peer-reviewed. The authors evaluated an event matcher, reviewed 2,738 findings from two models, and ranked six LLM extractors and two human annotators. Manual review confirmed 89.4% and 88.6% of findings. The authors say GAVEL supports report-based comparison and revision without treating either timeline as ground truth.