Why Normalizing Flows Fail to Detect Out-of-Distribution Data

Why Normalizing Flows Fail to Detect Out-of-Distribution Data
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
P. Kirichenko;Pavel Izmailov;A. Wilson
P. Kirichenko;Pavel Izmailov;A. Wilson
中科院分区:
其他
文献类型:
--
作者:
P. Kirichenko;Pavel Izmailov;A. Wilson

文献摘要

被引文献

相似文献

检测分布外(OOD)数据对于鲁棒的机器学习系统至关重要。规范化流是灵活的深度生成模型,经常令人惊讶地无法区分分布内和分布外的数据:在服装图片上训练的流将更高的可能性分配给手写数字。我们研究了为什么正规化流在OOD检测中表现不佳。我们证明了流学习局部像素相关性和通用图像到潜在空间的转换,这些转换并不特定于目标图像数据集。我们表明,通过修改流耦合层的架构,我们可以使流偏向于学习目标数据的语义结构,从而提高OOD检测。我们的研究表明,使流体产生高保真图像的特性可能对OOD检测产生不利影响。
Detecting out-of-distribution (OOD) data is crucial for robust machine learning systems. Normalizing flows are flexible deep generative models that often surprisingly fail to distinguish between in- and out-of-distribution data: a flow trained on pictures of clothing assigns higher likelihood to handwritten digits. We investigate why normalizing flows perform poorly for OOD detection. We demonstrate that flows learn local pixel correlations and generic image-to-latent-space transformations which are not specific to the target image dataset. We show that by modifying the architecture of flow coupling layers we can bias the flow towards learning the semantic structure of the target data, improving OOD detection. Our investigation reveals that properties that enable flows to generate high-fidelity images can have a detrimental effect on OOD detection.