The potential and perils of preprocessing: Building new foundations

The potential and perils of preprocessing: Building new foundations
复制标题

预处理的潜力和危险:建立新的基础

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
X. Meng
X. Meng
中科院分区:
--
文献类型:
--
作者:
A. Blocker;X. Meng

文献摘要

被引文献

相似文献

预处理构成了广泛的统计和科学分析的一个经常被忽视的基础。然而,它充满了微妙和陷阱。在预处理中做出的决策约束了所有后续的分析,并且通常是不可逆的。因此,数据分析成为涉及数据收集、预处理和整理以及下游推理的所有各方的协作努力。即使每一方在提供给他们的信息和资源的情况下都尽了最大的努力,最终的结果仍然可能在传统的单相推理框架中达不到最好的结果。在我们进入“大数据”时代之际,这一点尤为重要。驱动这种数据爆炸的技术受制于复杂的新形式的测量误差。同时,我们正在积累越来越庞大的科学分析数据库。因此,预处理变得比以往任何时候都更加重要(也可能更加危险)。我们提出了一个多相推理下的预处理分析理论框架。我们为这一领域提供了一些初步的理论基础,包括分布式预处理,建立在先前的多重插值工作的基础上。我们以生物学和天体物理学中的两个问题来激励这个基础,说明多相陷阱和潜在的解决方案。这些例子也强调了多相分析背后的动机——实践和理论。我们证明,在某些情况下,多相推理在效率和鲁棒性方面甚至可以超过标准的单相估计。我们的工作为进一步研究预处理背后的统计原理提供了几个丰富的途径。为了处理日益复杂和庞大的数据,我们必须确保我们的推论建立在可靠的输入和合理的原则之上。因此,预处理的原则性研究是统计研究的一个重要方向。重新参数化的作用,(3)渐近区域和多相效率,以及(4)多相推理中的鲁棒性问题。
Preprocessing forms an oft-neglected foundation for a wide range of statistical and scientific analyses. How-ever, it is rife with subtleties and pitfalls. Decisions made in preprocessing constrain all later analyses and are typically irreversible. Hence, data analysis becomes a collaborative endeavor by all parties involved in data collection, preprocessing and curation, and downstream inference. Even if each party has done its best given the information and resources available to them, the final result may still fall short of the best possible in the traditional single-phase inference framework. This is particularly relevant as we enter the era of “big data”. The technologies driving this data explosion are subject to complex new forms of measurement error. Simultaneously, we are accumulating increasingly massive databases of scientific analyses. As a result, preprocessing has become more vital (and potentially more dangerous) than ever before. We propose a theoretical framework for the analysis of preprocessing under the banner of multiphase inference. We provide some initial theoretical foundations for this area, including distributed preprocessing, building upon previous work in multiple imputation. We motivate this foundation with two problems from biology and astrophysics, illustrating multiphase pitfalls and potential solutions. These examples also emphasize the motivations behind multiphase analyses—both practical and theoretical. We demonstrate that multiphase inferences can, in some cases, even surpass standard single-phase estimators in efficiency and robustness. Our work suggests several rich paths for further research into the statistical principles underlying preprocessing. To tackle our increasingly complex and massive data, we must ensure that our inferences are built upon solid inputs and sound principles. Principled investigation of preprocessing is thus a vital direction for statistical research. the role of reparameterization, (3) asymptotic regimes and multiphase efficiency, and (4) the issue of robustness in multiphase inference.