Removing the influence of group variables in high-dimensional predictive modelling.

Removing the influence of group variables in high-dimensional predictive modelling.
复制标题

DOI:
10.1111/rssa.12613
复制
发表时间:
2021-07
影响因子:
2
通讯作者:
Dunson, David B.
Dunson, David B.
中科院分区:
数学4区
文献类型:
--
作者:
Aliverti, Emanuele;Lum, Kristian;Johndrow, James E.;Dunson, David B.

文献摘要

参考文献

相似文献

在许多应用领域,预测模型用于支持或做出重要决策。人们越来越意识到,这些模型可能包含虚假的或其他不可取的相关性。这种相关性可能来自各种来源,包括批次效应、系统测量误差或抽样偏差。如果没有明确的调整,使用这些数据训练的机器学习算法可能会产生糟糕的样本外预测,从而传播这些不期望的相关性。我们提出了一种方法来预处理的训练数据,产生一个调整后的数据集,是统计独立的滋扰变量与最小的信息损失。我们开发了一种概念上简单的方法,用于在高维设置中基于约束形式的矩阵分解创建调整后的数据集。由此产生的数据集可以用于任何预测算法,保证预测在统计上独立于组变量。我们开发了一个可扩展的算法实现的方法,沿着理论支持的形式的独立性保证和最优性。该方法说明了一些模拟的例子,并适用于两个案例研究:从大脑扫描数据中删除机器特定的相关性,并删除用于预测累犯的数据集的种族和民族信息。在这两个应用程序中,去除不期望的相关性的动机是完全不同的,这说明了我们的方法的广泛适用性。
In many application areas, predictive models are used to support or make important decisions. There is increasing awareness that these models may contain spurious or otherwise undesirable correlations. Such correlations may arise from a variety of sources, including batch effects, systematic measurement errors, or sampling bias. Without explicit adjustment, machine learning algorithms trained using these data can produce poor out-of-sample predictions which propagate these undesirable correlations. We propose a method to pre-process the training data, producing an adjusted dataset that is statistically independent of the nuisance variables with minimum information loss. We develop a conceptually simple approach for creating an adjusted dataset in high-dimensional settings based on a constrained form of matrix decomposition. The resulting dataset can then be used in any predictive algorithm with the guarantee that predictions will be statistically independent of the group variable. We develop a scalable algorithm for implementing the method, along with theory support in the form of independence guarantees and optimality. The method is illustrated on some simulation examples and applied to two case studies: removing machine-specific correlations from brain scan data, and removing race and ethnicity information from a dataset used to predict recidivism. That the motivation for removing undesirable correlations is quite different in the two applications illustrates the broad applicability of our approach.
DOI: 10.1093/bioinformatics/btg385
发表时间: 2004-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Benito, M;Parker, J;Marron, JS
通讯作者: Marron, JS
DOI: 10.1006/nimg.2001.1037
发表时间: 2002-04-01
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Genovese, CR;Lazar, NA;Nichols, T
通讯作者: Nichols, T
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D
DOI: 10.1016/j.spl.2018.02.028
发表时间: 2018-05-01
影响因子: 0.8
作者:
Dunson, David B.
通讯作者: Dunson, David B.
DOI: 10.1038/nn.4361
发表时间: 2016-08-26
影响因子: 25
作者:
Glasser MF;Smith SM;Marcus DS;Andersson JL;Auerbach EJ;Behrens TE;Coalson TS;Harms MP;Jenkinson M;Moeller S;Robinson EC;Sotiropoulos SN;Xu J;Yacoub E;Ugurbil K;Van Essen DC
通讯作者: Van Essen DC