Methods for correcting inference based on outcomes predicted by machine learning.

Methods for correcting inference based on outcomes predicted by machine learning.
复制标题

DOI:
10.1073/pnas.2001238117
复制
发表时间:
2020-12-01
影响因子:
11.1
通讯作者:
Leek JT
Leek JT
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Wang S;McCormick TH;Leek JT

文献摘要

参考文献

被引文献

相似文献

机器学习现在已经被应用于整个科学领域。研究人员通常在下游统计分析中使用随机森林或深度神经网络的预测,就好像它们是观察数据一样。我们表明,这种方法可能会导致极端的偏见和不受控制的方差在下游统计模型。我们提出了一个统计调整,以纠正有偏见的推理回归模型使用预测的结果,无论机器学习模型用于作出这些预测。医学和公共卫生领域的许多现代问题都利用机器学习方法来预测基于可观察协变量的结果。在许多情况下,预测结果用于随后的统计分析,通常不考虑观察结果和预测结果之间的区别。我们把带有预测结果的推理称为后预测推理。在本文中,我们开发了使用任意复杂的机器学习模型(包括随机森林和深度神经网络)预测的结果来纠正统计推断的方法。我们没有试图从每个机器学习算法的第一原理中得出校正,而是观察到观察结果和预测结果之间的关系通常存在低维且易于建模的表示。我们构建了一种用于后预测推理的方法,该方法自然适合标准机器学习框架,其中数据被划分为训练,测试和验证集。我们在训练集中训练预测模型,估计测试集中观察结果和预测结果之间的关系,并使用该关系来纠正验证集中的后续推断。我们证明了我们的后预测推理(postpi)方法可以纠正偏差,提高方差估计和随后的统计推断与预测结果。为了展示我们的方法的广泛适用性,我们展示了postpi可以在两个不同的领域改进推理:在重新利用的基因表达数据中建模预测的表型和在死因推断数据中建模预测的死亡原因。我们的方法可以通过一个开源的R包获得:https://github.com/leekgroup/postpi。
Machine learning is now being used across the entire scientific enterprise. Researchers commonly use the predictions from random forests or deep neural networks in downstream statistical analysis as if they were observed data. We show that this approach can lead to extreme bias and uncontrolled variance in downstream statistical models. We propose a statistical adjustment to correct biased inference in regression models using predicted outcomes—regardless of the machine-learning model used to make those predictions. Many modern problems in medicine and public health leverage machine-learning methods to predict outcomes based on observable covariates. In a wide array of settings, predicted outcomes are used in subsequent statistical analysis, often without accounting for the distinction between observed and predicted outcomes. We call inference with predicted outcomes postprediction inference. In this paper, we develop methods for correcting statistical inference using outcomes predicted with arbitrarily complicated machine-learning models including random forests and deep neural nets. Rather than trying to derive the correction from first principles for each machine-learning algorithm, we observe that there is typically a low-dimensional and easily modeled representation of the relationship between the observed and predicted outcomes. We build an approach for postprediction inference that naturally fits into the standard machine-learning framework where the data are divided into training, testing, and validation sets. We train the prediction model in the training set, estimate the relationship between the observed and predicted outcomes in the testing set, and use that relationship to correct subsequent inference in the validation set. We show our postprediction inference (postpi) approach can correct bias and improve variance estimation and subsequent statistical inference with predicted outcomes. To show the broad range of applicability of our approach, we show postpi can improve inference in two distinct fields: modeling predicted phenotypes in repurposed gene expression data and modeling predicted causes of death in verbal autopsy data. Our method is available through an open-source R package: https://github.com/leekgroup/postpi.
DOI: 10.1016/j.amepre.2015.08.031
发表时间: 2016-03
影响因子: 5.5
作者:
Khoury MJ;Iademarco MF;Riley WT
通讯作者: Riley WT
遗传对人体组织基因表达的影响。
DOI: 10.1038/nature24277
发表时间: 2017-10-11
期刊: Nature
影响因子: 64.8
作者:
GTEx Consortium;Laboratory, Data Analysis &Coordinating Center (LDACC)—Analysis Working Group;Statistical Methods groups—Analysis Working Group;Enhancing GTEx (eGTEx) groups;NIH Common Fund;NIH/NCI;NIH/NHGRI;NIH/NIMH;NIH/NIDA;Biospecimen Collection Source Site—NDRI;Biospecimen Collection Source Site—RPCI;Biospecimen Core Resource—VARI;Brain Bank Repository—University of Miami Brain Endowment Bank;Leidos Biomedical—Project Management;ELSI Study;Genome Browser Data Integration &Visualization—EBI;Genome Browser Data Integration &Visualization—UCSC Genomics Institute, University of California Santa Cruz;Lead analysts:;Laboratory, Data Analysis &Coordinating Center (LDACC):;NIH program management:;Biospecimen collection:;Pathology:;eQTL manuscript working group:;Battle A;Brown CD;Engelhardt BE;Montgomery SB
通讯作者: Montgomery SB
DOI: 10.1093/nar/gkw852
发表时间: 2017-01-25
影响因子: 14.9
作者:
Collado-Torres L;Nellore A;Frazee AC;Wilks C;Love MI;Langmead B;Irizarry RA;Leek JT;Jaffe AE
通讯作者: Jaffe AE
DOI: 10.1101/gr.155192.113
发表时间: 2014-01
期刊: Genome research
影响因子: 7
作者:
Battle A;Mostafavi S;Zhu X;Potash JB;Weissman MM;McCormick C;Haudenschild CD;Beckman KB;Shi J;Mei R;Urban AE;Montgomery SB;Levinson DF;Koller D
通讯作者: Koller D
DOI: 10.1038/nature09944
发表时间: 2011-05-12
期刊: NATURE
影响因子: 64.8
作者:
Arumugam, Manimozhiyan;Raes, Jeroen;Pelletier, Eric;Le Paslier, Denis;Yamada, Takuji;Mende, Daniel R.;Fernandes, Gabriel R.;Tap, Julien;Bruls, Thomas;Batto, Jean-Michel;Bertalan, Marcelo;Borruel, Natalia;Casellas, Francesc;Fernandez, Leyden;Gautier, Laurent;Hansen, Torben;Hattori, Masahira;Hayashi, Tetsuya;Kleerebezem, Michiel;Kurokawa, Ken;Leclerc, Marion;Levenez, Florence;Manichanh, Chaysavanh;Nielsen, H. Bjorn;Nielsen, Trine;Pons, Nicolas;Poulain, Julie;Qin, Junjie;Sicheritz-Ponten, Thomas;Tims, Sebastian;Torrents, David;Ugarte, Edgardo;Zoetendal, Erwin G.;Wang, Jun;Guarner, Francisco;Pedersen, Oluf;de Vos, Willem M.;Brunak, Soren;Dore, Joel;Weissenbach, Jean;Ehrlich, S. Dusko;Bork, Peer
通讯作者: Bork, Peer