An Approach to Identifying and Quantifying Bias in Biomedical Data

An Approach to Identifying and Quantifying Bias in Biomedical Data
复制标题

DOI:
10.1142/9789811270611_0029
复制
发表时间:
2022-11
影响因子:
--
通讯作者:
M. C. De Paolis Kaluza;Shantanu Jain;P. Radivojac
M. C. De Paolis Kaluza;Shantanu Jain;P. Radivojac
中科院分区:
--
文献类型:
--
作者:
M. C. De Paolis Kaluza;Shantanu Jain;P. Radivojac

文献摘要

相似文献

众所周知,数据偏差是开发值得信赖的机器学习模型及其应用于许多生物医学问题的障碍。当有偏数据被怀疑时,必须放松标记数据代表总体的假设,并且必须开发利用典型代表性未标记数据的方法。为了减轻不具有代表性的数据的不利影响,我们考虑了一个二进制半监督设置,并专注于识别标记的数据是否有偏见,以及在多大程度上。我们假设类条件分布是由一系列在标记和未标记数据中以不同比例表示的分量分布生成的。我们还假设训练数据可以转换为多元高斯分布的嵌套混合,并随后由其建模。然后,我们开发了一个多样本期望最大化算法,从组合数据中学习模型的所有个人和共享参数。使用这些参数,我们开发了一个统计测试的一般形式的偏差在标记数据的存在,并估计这种偏差的水平,通过计算相应的类条件分布之间的距离标记和未标记的数据。我们首先研究了合成数据的新方法,以了解它们的行为,然后将它们应用于现实世界的生物医学数据,以提供证据证明偏差估计程序是可能的和有效的。
Data biases are a known impediment to the development of trustworthy machine learning models and their application to many biomedical problems. When biased data is suspected, the assumption that the labeled data is representative of the population must be relaxed and methods that exploit a typically representative unlabeled data must be developed. To mitigate the adverse effects of unrepresentative data, we consider a binary semi-supervised setting and focus on identifying whether the labeled data is biased and to what extent. We assume that the class-conditional distributions were generated by a family of component distributions represented at different proportions in labeled and unlabeled data. We also assume that the training data can be transformed to and subsequently modeled by a nested mixture of multivariate Gaussian distributions. We then develop a multi-sample expectation-maximization algorithm that learns all individual and shared parameters of the model from the combined data. Using these parameters, we develop a statistical test for the presence of the general form of bias in labeled data and estimate the level of this bias by computing the distance between corresponding class-conditional distributions in labeled and unlabeled data. We first study the new methods on synthetic data to understand their behavior and then apply them to real-world biomedical data to provide evidence that the bias estimation procedure is both possible and effective.