推论
要价
无效假设
序列(生物学)
差异(会计)
计算机科学
计量经济学
有界函数
统计假设检验
统计推断
置信区间
替代假设
停止规则
数学
统计
人工智能
经济
数学优化
数学分析
经济
会计
生物
遗传学
作者
Yo Joong Choe,Aaditya Ramdas
标识
DOI:10.48550/arxiv.2110.00115
摘要
Consider two forecasters, each making a single prediction for a sequence of events over time. We ask a relatively basic question: how might we compare these forecasters, either online or post-hoc, while avoiding unverifiable assumptions on how the forecasts and outcomes were generated? In this paper, we present a rigorous answer to this question by designing novel sequential inference procedures for estimating the time-varying difference in forecast scores. To do this, we employ confidence sequences (CS), which are sequences of confidence intervals that can be continuously monitored and are valid at arbitrary data-dependent stopping times ("anytime-valid"). The widths of our CSs are adaptive to the underlying variance of the score differences. Underlying their construction is a game-theoretic statistical framework, in which we further identify e-processes and p-processes for sequentially testing a weak null hypothesis -- whether one forecaster outperforms another on average (rather than always). Our methods do not make distributional assumptions on the forecasts or outcomes; our main theorems apply to any bounded scores, and we later provide alternative methods for unbounded scores. We empirically validate our approaches by comparing real-world baseball and weather forecasters.
科研通智能强力驱动
Strongly Powered by AbleSci AI