Eliminating accidental deviations to minimize generalization error and maximize replicability: Applications in connectomics and genomics.

Eliminating accidental deviations to minimize generalization error and maximize replicability: Applications in connectomics and genomics.
复制标题

消除偶然偏差以最小化泛化错误和最大化可复制性:在连接学和基因组学中的应用。

DOI:
10.1371/journal.pcbi.1009279
复制
发表时间:
2021-09
影响因子:
4.3
通讯作者:
Vogelstein JT
Vogelstein JT
中科院分区:
生物学2区
文献类型:
--
作者:
Bridgeford EW;Wang S;Wang Z;Xu T;Craddock C;Dey J;Kiar G;Gray-Roncal W;Colantuoni C;Douville C;Noble S;Priebe CE;Caffo B;Milham M;Zuo XN;Consortium for Reliability and Reproducibility;Vogelstein JT

文献摘要

参考文献

被引文献

相似文献

可复制性,即复制科学发现的能力,是科学发现和临床实用的先决条件。令人不安的是,我们正处于一场可复制危机之中。可重复性的一个关键是同一项目(例如,实验样本或临床参与者)在固定实验约束下的多个测量彼此相对相似。因此,量化偶然偏差--如测量误差--相对于系统性偏差--如个体差异--的相对贡献的统计数据至关重要。我们证明,现有的可复制性统计,如类内相关系数和指纹,在非常简单的设置下无法充分区分意外偏差和系统性偏差。因此,我们提出了一种新的统计量,可区分性,它量化了一个人的样本彼此相对相似的程度,而不限制数据为单变量、高斯甚至欧几里德。利用这个统计量,我们介绍了通过提高可区分性来优化实验设计的可能性,并证明了优化可区分性可以改善后续推理任务的性能界限。在大量的模拟和真实数据集中(侧重于脑成像和基因组学演示),只有优化数据可区分性才能提高每个数据集的所有后续推理任务的性能。因此,我们认为,设计实验和分析以优化区分度可能是解决可复制性危机的关键一步,更广泛地说,是减少意外测量误差的关键一步。近几十年来,数据的大小和复杂性呈指数级增长。不幸的是,现代数据集规模的扩大带来了许多新的挑战。目前,我们正处于可复制性危机之中,科学发现无法复制到新的数据集。我们认为,测量程序和测量处理管道中的困难,再加上复杂的高分辨率测量的涌入,是可复制危机的核心。如果测量本身是不可复制的,我们还能有什么希望将测量结果用于可复制的科学发现呢?我们引入了“可区分性”统计量,它量化了测量之间的可区分性,对基本测量的结构没有限制。我们证明,可区分的策略往往是在下游科学问题上提供更好准确性的策略。在这一背景下,我们在来自神经成像和基因组学的两个不同的数据集上展示了可识别性相对于竞争方法的实用性。总之,我们认为这些结果表明了设计实验方案和分析程序的价值,这些程序优化了可区分性。
Replicability, the ability to replicate scientific findings, is a prerequisite for scientific discovery and clinical utility. Troublingly, we are in the midst of a replicability crisis. A key to replicability is that multiple measurements of the same item (e.g., experimental sample or clinical participant) under fixed experimental constraints are relatively similar to one another. Thus, statistics that quantify the relative contributions of accidental deviations—such as measurement error—as compared to systematic deviations—such as individual differences—are critical. We demonstrate that existing replicability statistics, such as intra-class correlation coefficient and fingerprinting, fail to adequately differentiate between accidental and systematic deviations in very simple settings. We therefore propose a novel statistic, discriminability, which quantifies the degree to which an individual’s samples are relatively similar to one another, without restricting the data to be univariate, Gaussian, or even Euclidean. Using this statistic, we introduce the possibility of optimizing experimental design via increasing discriminability and prove that optimizing discriminability improves performance bounds in subsequent inference tasks. In extensive simulated and real datasets (focusing on brain imaging and demonstrating on genomics), only optimizing data discriminability improves performance on all subsequent inference tasks for each dataset. We therefore suggest that designing experiments and analyses to optimize discriminability may be a crucial step in solving the replicability crisis, and more generally, mitigating accidental measurement error. In recent decades, the size and complexity of data has grown exponentially. Unfortunately, the increased scale of modern datasets brings many new challenges. At present, we are in the midst of a replicability crisis, in which scientific discoveries fail to replicate to new datasets. Difficulties in the measurement procedure and measurement processing pipelines coupled with the influx of complex high-resolution measurements, we believe, are at the core of the replicability crisis. If measurements themselves are not replicable, what hope can we have that we will be able to use the measurements for replicable scientific findings? We introduce the “discriminability” statistic, which quantifies how discriminable measurements are from one another, without limitations on the structure of the underlying measurements. We prove that discriminable strategies tend to be strategies which provide better accuracy on downstream scientific questions. We demonstrate the utility of discriminability over competing approaches in this context on two disparate datasets from both neuroimaging and genomics. Together, we believe these results suggest the value of designing experimental protocols and analysis procedures which optimize the discriminability.
DOI: 10.1126/scitranslmed.aaf5027
发表时间: 2016-06-01
影响因子: 17.1
作者:
Goodman, Steven N.;Fanelli, Daniele;Ioannidis, John P. A.
通讯作者: Ioannidis, John P. A.
DOI: 10.2307/2092790
发表时间: 1969-01-01
影响因子: 9.1
作者:
HEISE, DR
通讯作者: HEISE, DR
DOI: 10.1371/journal.pmed.0020124
发表时间: 2005-08-01
期刊: PLOS MEDICINE
影响因子: 15.8
作者:
Ioannidis, JPA
通讯作者: Ioannidis, JPA
探索人类大脑功能的科学
DOI: 10.1073/pnas.0911855107
发表时间: 2010-03-09
影响因子: 11.1
作者:
Biswal, Bharat B.;Mennes, Maarten;Milham, Michael P.
通讯作者: Milham, Michael P.
DOI: 10.1016/j.neuroimage.2017.02.036
发表时间: 2017-04-15
期刊: NeuroImage
影响因子: 5.7
作者:
Liu TT;Nalci A;Falahpour M
通讯作者: Falahpour M