TESTING FOR OUTLIERS WITH CONFORMAL P-VALUES

TESTING FOR OUTLIERS WITH CONFORMAL P-VALUES
复制标题

DOI:
10.1214/22-aos2244
复制
发表时间:
2023-02-01
影响因子:
4.5
通讯作者:
Sesia, Matteo
Sesia, Matteo
中科院分区:
数学1区
文献类型:
--
作者:
Bates, Stephen;Candes, Emmanuel;Sesia, Matteo

文献摘要

被引文献

相似文献

本文从多重检验的角度研究了非参数异常值检测中p值的构造。目标是测试新的独立样本是否属于与参考数据集相同的分布或离群值。我们提出了一个基于共形推理的解决方案,这是一个通用框架,可以产生勉强有效但相互依赖的p值,用于不同的测试点。我们证明了这些p值是正相关的,并使准确的错误发现率控制,虽然在一个相对较弱的边缘意义。然后,我们引入了一种新的方法来计算p值,这些p值在训练数据上是有条件有效的,并且对于不同的测试点彼此独立;这为更强的I类错误保证铺平了道路。我们的结果偏离经典的共形推理,因为我们利用浓度不等式,而不是组合参数,以建立我们的有限样本保证。此外,我们的技术还产生了一个统一的置信界的假阳性率的任何离群值检测算法,作为应用于其原始统计数据的阈值的函数。最后,通过对真实的和模拟数据的实验,证明了我们的结果的相关性。
This paper studies the construction of p-values for nonparametric out-lier detection, from a multiple-testing perspective. The goal is to test whether new independent samples belong to the same distribution as a reference data set or are outliers. We propose a solution based on conformal inference, a general framework yielding p-values that are marginally valid but mutually dependent for different test points. We prove these p-values are positively de-pendent and enable exact false discovery rate control, although in a relatively weak marginal sense. We then introduce a new method to compute p-values that are valid conditionally on the training data and independent of each other for different test points; this paves the way to stronger type-I error guarantees. Our results depart from classical conformal inference as we leverage con-centration inequalities rather than combinatorial arguments to establish our finite-sample guarantees. Further, our techniques also yield a uniform confi-dence bound for the false positive rate of any outlier detection algorithm, as a function of the threshold applied to its raw statistics. Finally, the relevance of our results is demonstrated by experiments on real and simulated data.