How to remove or control confounds in predictive models, with applications to brain biomarkers.

How to remove or control confounds in predictive models, with applications to brain biomarkers.
复制标题

DOI:
10.1093/gigascience/giac014
复制
发表时间:
2022-03-12
期刊:
影响因子:
9.2
通讯作者:
Thirion B
Thirion B
中科院分区:
生物学2区
文献类型:
--
作者:
Chyzhyk D;Varoquaux G;Milham M;Thirion B

文献摘要

参考文献

被引文献

相似文献

随着数据量的增加和更容易获得的计算方法,神经科学越来越依赖于机器学习的预测建模,例如提取疾病生物标志物。然而,一个成功的预测可能会捕捉到与结果相关的混淆效应,而不是特定于感兴趣的结果的大脑特征。例如,由于患者在扫描仪中往往比对照组移动更多,疾病状况的成像生物标志物可能主要反映头部运动,导致资源使用效率低下和对生物标志物的错误解释。在这里,我们研究如何使控制混杂的统计方法适应预测建模设置。我们回顾了如何训练不受这种虚假效应驱动的预测器。我们还展示了如何在混杂数据集的基础上测量这些生物标志物的无偏预测准确性。为此目的,必须修改交叉验证以考虑妨害效应。为了指导理解和实际建议,我们应用各种策略来评估模拟数据和人群脑成像设置中存在混淆的预测模型。理论和实证研究表明,反发现不应该同时应用于训练数据和测试数据:只对训练数据建模混淆的影响,而应该与去除混淆分离。交叉验证分离了讨厌的影响,提供了额外的信息:无混杂预测的准确性。
With increasing data sizes and more easily available computational methods, neurosciences rely more and more on predictive modeling with machine learning, e.g., to extract disease biomarkers. Yet, a successful prediction may capture a confounding effect correlated with the outcome instead of brain features specific to the outcome of interest. For instance, because patients tend to move more in the scanner than controls, imaging biomarkers of a disease condition may mostly reflect head motion, leading to inefficient use of resources and wrong interpretation of the biomarkers. Here we study how to adapt statistical methods that control for confounds to predictive modeling settings. We review how to train predictors that are not driven by such spurious effects. We also show how to measure the unbiased predictive accuracy of these biomarkers, based on a confounded dataset. For this purpose, cross-validation must be modified to account for the nuisance effect. To guide understanding and practical recommendations, we apply various strategies to assess predictive models in the presence of confounds on simulated data and population brain imaging settings. Theoretical and empirical studies show that deconfounding should not be applied to the train and test data jointly: modeling the effect of confounds, on the training data only, should instead be decoupled from removing confounds. Cross-validation that isolates nuisance effects gives an additional piece of information: confound-free prediction accuracy.
DOI: 10.3389/fninf.2011.00013
发表时间: 2011
影响因子: 3.5
作者:
Gorgolewski K;Burns CD;Madison C;Clark D;Halchenko YO;Waskom ML;Ghosh SS
通讯作者: Ghosh SS
DOI: 10.3389/fninf.2014.00014
发表时间: 2014
影响因子: 3.5
作者:
Abraham A;Pedregosa F;Eickenberg M;Gervais P;Mueller A;Kossaifi J;Gramfort A;Thirion B;Varoquaux G
通讯作者: Varoquaux G
DOI: 10.1002/hbm.23653
发表时间: 2017-08
影响因子: 4.8
作者:
Geerligs L;Tsvetanov KA;Cam-Can;Henson RN
通讯作者: Henson RN
DOI: 10.1016/j.neuroimage.2010.01.005
发表时间: 2010-04-15
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Franke, Katja;Ziegler, Gabriel;Gaser, Christian
通讯作者: Gaser, Christian
DOI: 10.1002/sim.1657
发表时间: 2004-03-15
影响因子: 2
作者:
Brumback, BA;Hernán, MA;Robins, JM
通讯作者: Robins, JM