Principled OOD Detection via Multiple Testing

Principled OOD Detection via Multiple Testing
复制标题

DOI:
10.1109/isit54713.2023.10206581
复制
发表时间:
2023-06
期刊:
2023 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
A. Magesh;V. Veeravalli;Anirban Roy;Susmit Jha
A. Magesh;V. Veeravalli;Anirban Roy;Susmit Jha
中科院分区:
其他
文献类型:
--
作者:
A. Magesh;V. Veeravalli;Anirban Roy;Susmit Jha

文献摘要

相似文献

我们研究了分布外(OOD)检测问题,即在推理时检测机器学习(ML)模型的输出是否可信。虽然在以前的工作中已经提出了一些OOD检测的测试,研究这个问题的正式框架是缺乏的。我们提出了OOD概念的定义,包括输入分布和ML模型,这为构建强大的OOD检测测试提供了见解。我们还提出了一个多假设检验的启发程序,系统地结合联合收割机的任何数量的不同的统计数据,从ML模型使用保形p值。我们进一步提供了强有力的保证,错误地将分布样本分类为OOD的概率。在我们的实验中,我们发现,在以前的工作中提出的基于阈值的测试在特定的设置中表现良好,但在不同的OOD实例中并不一致。相比之下,我们提出的结合多种统计数据的方法在不同的数据集和神经网络中表现一致。
We study the problem of Out-of-Distribution (OOD) detection, that is, detecting whether a Machine Learning (ML) model's output can be trusted at inference time. While a number of tests for OOD detection have been proposed in prior work, a formal framework for studying this problem is lacking. We propose a definition for the notion of OOD that includes both the input distribution and the ML model, which provides insights for the construction of powerful tests for OOD detection. We also propose a multiple hypothesis testing inspired procedure to systematically combine any number of different statistics from the ML model using conformal p-values. We further provide strong guarantees on the probability of incorrectly classifying an in-distribution sample as OOD. In our experiments, we find that threshold-based tests proposed in prior work perform well in specific settings, but not uniformly well across different OOD instances. In contrast, our proposed method that combines multiple statistics performs uniformly well across different datasets and neural networks.