Methods for correcting inference based on outcomes predicted by machine learning.
Methods for correcting inference based on outcomes predicted by machine learning.
复制标题
DOI:
10.1073/pnas.2001238117
复制
发表时间:
2020-12-01
影响因子:
11.1
通讯作者:
Leek JT
中科院分区:
文献类型:
--
作者:
Wang S;McCormick TH;Leek JT
Machine learning is now being used across the entire scientific enterprise. Researchers commonly use the predictions from random forests or deep neural networks in downstream statistical analysis as if they were observed data. We show that this approach can lead to extreme bias and uncontrolled variance in downstream statistical models. We propose a statistical adjustment to correct biased inference in regression models using predicted outcomes—regardless of the machine-learning model used to make those predictions. Many modern problems in medicine and public health leverage machine-learning methods to predict outcomes based on observable covariates. In a wide array of settings, predicted outcomes are used in subsequent statistical analysis, often without accounting for the distinction between observed and predicted outcomes. We call inference with predicted outcomes postprediction inference. In this paper, we develop methods for correcting statistical inference using outcomes predicted with arbitrarily complicated machine-learning models including random forests and deep neural nets. Rather than trying to derive the correction from first principles for each machine-learning algorithm, we observe that there is typically a low-dimensional and easily modeled representation of the relationship between the observed and predicted outcomes. We build an approach for postprediction inference that naturally fits into the standard machine-learning framework where the data are divided into training, testing, and validation sets. We train the prediction model in the training set, estimate the relationship between the observed and predicted outcomes in the testing set, and use that relationship to correct subsequent inference in the validation set. We show our postprediction inference (postpi) approach can correct bias and improve variance estimation and subsequent statistical inference with predicted outcomes. To show the broad range of applicability of our approach, we show postpi can improve inference in two distinct fields: modeling predicted phenotypes in repurposed gene expression data and modeling predicted causes of death in verbal autopsy data. Our method is available through an open-source R package: https://github.com/leekgroup/postpi.
登录
查看更多内容
影响因子:
5.5
作者:
Khoury MJ;Iademarco MF;Riley WT
通讯作者:
Riley WT
影响因子:
64.8
作者:
GTEx Consortium;Laboratory, Data Analysis &Coordinating Center (LDACC)—Analysis Working Group;Statistical Methods groups—Analysis Working Group;Enhancing GTEx (eGTEx) groups;NIH Common Fund;NIH/NCI;NIH/NHGRI;NIH/NIMH;NIH/NIDA;Biospecimen Collection Source Site—NDRI;Biospecimen Collection Source Site—RPCI;Biospecimen Core Resource—VARI;Brain Bank Repository—University of Miami Brain Endowment Bank;Leidos Biomedical—Project Management;ELSI Study;Genome Browser Data Integration &Visualization—EBI;Genome Browser Data Integration &Visualization—UCSC Genomics Institute, University of California Santa Cruz;Lead analysts:;Laboratory, Data Analysis &Coordinating Center (LDACC):;NIH program management:;Biospecimen collection:;Pathology:;eQTL manuscript working group:;Battle A;Brown CD;Engelhardt BE;Montgomery SB
通讯作者:
Montgomery SB
影响因子:
14.9
作者:
Collado-Torres L;Nellore A;Frazee AC;Wilks C;Love MI;Langmead B;Irizarry RA;Leek JT;Jaffe AE
通讯作者:
Jaffe AE
影响因子:
7
作者:
Battle A;Mostafavi S;Zhu X;Potash JB;Weissman MM;McCormick C;Haudenschild CD;Beckman KB;Shi J;Mei R;Urban AE;Montgomery SB;Levinson DF;Koller D
通讯作者:
Koller D
影响因子:
64.8
作者:
Arumugam, Manimozhiyan;Raes, Jeroen;Pelletier, Eric;Le Paslier, Denis;Yamada, Takuji;Mende, Daniel R.;Fernandes, Gabriel R.;Tap, Julien;Bruls, Thomas;Batto, Jean-Michel;Bertalan, Marcelo;Borruel, Natalia;Casellas, Francesc;Fernandez, Leyden;Gautier, Laurent;Hansen, Torben;Hattori, Masahira;Hayashi, Tetsuya;Kleerebezem, Michiel;Kurokawa, Ken;Leclerc, Marion;Levenez, Florence;Manichanh, Chaysavanh;Nielsen, H. Bjorn;Nielsen, Trine;Pons, Nicolas;Poulain, Julie;Qin, Junjie;Sicheritz-Ponten, Thomas;Tims, Sebastian;Torrents, David;Ugarte, Edgardo;Zoetendal, Erwin G.;Wang, Jun;Guarner, Francisco;Pedersen, Oluf;de Vos, Willem M.;Brunak, Soren;Dore, Joel;Weissenbach, Jean;Ehrlich, S. Dusko;Bork, Peer
通讯作者:
Bork, Peer