The paper introduces a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, it improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The useful detail for practitioners: outcome-estimated competence supports better abstention decisions, while verbal confidence does not reliably identify when the model outperforms the external forecast. However, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market.
Practical AI/Report
New method explains self-driving car AI decisions in real time
Researchers developed CW-Net, a method that translates an autonomous vehicle AI planner’s reasoning into understandable concepts and outputs explanations alongside its trajectory. Road tests…
