Bayesian posteriors for arbitrarily rare events

Bayesian posteriors for arbitrarily rare events
复制标题

DOI:
10.1073/pnas.1618780114
复制
发表时间:
2016-08
期刊:
Proceedings of the National Academy of Sciences
影响因子:
--
通讯作者:
D. Fudenberg;Kevin He;L. Imhof
D. Fudenberg;Kevin He;L. Imhof
中科院分区:
其他
文献类型:
--
作者:
D. Fudenberg;Kevin He;L. Imhof

文献摘要

被引文献

相似文献

从药物安全测试到博弈论学习模型的许多决策问题都需要在两个事件的可能性之间进行贝叶斯比较。当这两个事件都非常罕见时,需要一个大的数据集来获得高概率的正确决策。在以前的工作中,最好的结果要求数据量增长得如此之快,以至于对罕见事件的观测数量的期望达到爆炸式增长。我们证明,对于一个大的先验类,这个期望超过一个先验相关常数就足够了。然而,如果没有对先验的一些限制,结果就会失败,并且我们对数据大小的条件是最弱的。我们研究贝叶斯观察者需要多少数据才能正确推断两个事件的相对可能性,当两个事件都是任意罕见的。每一阶段,掷出一个蓝骰子或一个红骰子。两个骰子落在第1面,概率是未知的p1和q1,可以任意低。给定一个p1≥cq1的数据生成过程,我们感兴趣的是需要多少数据才能保证观察者p1的贝叶斯后验均值在高概率下超过q1的贝叶斯后验均值(1−δ)c倍。如果两个骰子的先验密度在参数空间内部是正的,并且在边界处表现得像幂函数,那么对于每一个λ > 0,存在一个有限的N,使得观察者在N个周期后得到这样的推断,无论np1≥N,概率至少为1−λ。n和p1的条件是最好的。如果其中一个先验密度在边界处以指数速度收敛于零,则结果可能失效。
Significance Many decision problems in contexts ranging from drug safety tests to game-theoretic learning models require Bayesian comparisons between the likelihoods of two events. When both events are arbitrarily rare, a large data set is needed to reach the correct decision with high probability. The best result in previous work requires the data size to grow so quickly with rarity that the expectation of the number of observations of the rare event explodes. We show for a large class of priors that it is enough that this expectation exceeds a prior-dependent constant. However, without some restrictions on the prior the result fails, and our condition on the data size is the weakest possible. We study how much data a Bayesian observer needs to correctly infer the relative likelihoods of two events when both events are arbitrarily rare. Each period, either a blue die or a red die is tossed. The two dice land on side 1 with unknown probabilities p1 and q1, which can be arbitrarily low. Given a data-generating process where p1≥cq1, we are interested in how much data are required to guarantee that with high probability the observer’s Bayesian posterior mean for p1 exceeds (1−δ)c times that for q1. If the prior densities for the two dice are positive on the interior of the parameter space and behave like power functions at the boundary, then for every ϵ> 0, there exists a finite N so that the observer obtains such an inference after n periods with probability at least 1−ϵ whenever np1≥N. The condition on n and p1 is the best possible. The result can fail if one of the prior densities converges to zero exponentially fast at the boundary.