How Much Does Your Data Exploration Overfit? Controlling Bias via Information Usage

How Much Does Your Data Exploration Overfit? Controlling Bias via Information Usage
复制标题

DOI:
10.1109/tit.2019.2945779
复制
发表时间:
2020-01-01
影响因子:
2.5
通讯作者:
Zou, James
Zou, James
中科院分区:
计算机科学2区
文献类型:
--
作者:
Russo, Daniel;Zou, James

文献摘要

被引文献

相似文献

现代数据是混乱且高维的,并且通常不清楚先验什么是正确的问题。相反,分析师通常需要使用数据来搜索要执行的有趣分析和要测试的假设。这是一个自适应过程,其中接下来要执行的分析的选择取决于先前对相同数据的分析结果。最终,报告的结果可能会受到数据的严重影响。人们普遍认为,这一过程即使是出于善意,也可能导致偏见和错误发现,从而导致科学可重复性危机。但是,尽管任何数据探索都会使标准统计理论无效,但经验表明,不同类型的探索性分析可能会导致不同程度的偏差,并且偏差的程度还取决于数据集的具体情况。在本文中,我们提出了一个通用信息使用框架来量化和可证明地限制任意探索性分析的偏差和其他错误指标。我们证明了基于互信息的界限在自然环境中是严格的,然后用它来严格洞察常用程序何时会或不会导致严重偏差的估计。通过信息使用的视角,我们分析了过滤、排名选择和聚类等特定探索过程的偏差。我们的总体框架也自然地激发了随机化技术,这些技术可以证明减少探索偏差,同时保留数据分析的实用性。我们讨论了我们的方法与差异隐私和盲数据分析的相关想法之间的联系,并用说明性模拟补充了我们的结果。
Modern data is messy and high-dimensional, and it is often not clear a priori what are the right questions to ask. Instead, the analyst typically needs to use the data to search for interesting analyses to perform and hypotheses to test. This is an adaptive process, where the choice of analysis to be performed next depends on the results of the previous analyses on the same data. Ultimately, which results are reported can be heavily influenced by the data. It is widely recognized that this process, even if well-intentioned, can lead to biases and false discoveries, contributing to the crisis of reproducibility in science. But while any data-exploration renders standard statistical theory invalid, experience suggests that different types of exploratory analysis can lead to disparate levels of bias, and the degree of bias also depends on the particulars of the data set. In this paper, we propose a general information usage framework to quantify and provably bound the bias and other error metrics of an arbitrary exploratory analysis. We prove that our mutual information based bound is tight in natural settings, and then use it to give rigorous insights into when commonly used procedures do or do not lead to substantially biased estimation. Through the lens of information usage, we analyze the bias of specific exploration procedures such as filtering, rank selection and clustering. Our general framework also naturally motivates randomization techniques that provably reduce exploration bias while preserving the utility of the data analysis. We discuss the connections between our approach and related ideas from differential privacy and blinded data analysis, and supplement our results with illustrative simulations.