Offspring GAN augments biased human genomic data

Offspring GAN augments biased human genomic data
复制标题

DOI:
10.1145/3535508.3545537
复制
发表时间:
2022-08
期刊:
Proceedings of the 13th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics
影响因子:
--
通讯作者:
Supratim Das;Xinghua Shi
Supratim Das;Xinghua Shi
中科院分区:
其他
文献类型:
--
作者:
Supratim Das;Xinghua Shi

文献摘要

被引文献

相似文献

长期以来,基因组数据一直被用于性状关联和疾病风险预测。近年来,许多这样的预测模型都是使用机器学习(ML)算法建立的。截至目前,人类基因组数据和其他生物医学数据在人们的种族方面受到抽样偏差的影响,因为大多数数据来自欧洲血统的人。对于其他人口组,较小的样本量可能会导致基于ML的预测模型对这些人口的预测结果不太理想。精确医学中对某些特定群体的次优预测可能会导致严重的后果,限制了该模型在现实世界问题中的适用性。由于为这些人群收集数据既耗时又昂贵,我们建议使用基于深度学习的模型来增强电子数据。现有的针对基因组数据的生成性对抗网络(GAN)模型,如种群规模基因组条件网络(PG-cGAN),可以在对相当无偏的数据进行训练时生成真实的基因组数据,但在对有偏数据进行训练时会失败,并遇到严重的模式崩溃。我们提出的模型--子代GAN,即使在强偏倚的基因组数据集中训练时,也可以解决模式崩溃问题。我们的结果证明了子代GAN产生现实的和多样化的标签感知数据的能力,这可以增加有限的真实数据来缓解基因组数据中的偏差和差异。我们还提出了一种基于子代GAN的隐私保护协议来保护基因组数据的隐私。
Genomic data have been used for trait association and disease risk prediction for a long time. In recent years, many such prediction models are built using machine learning (ML) algorithms. As of today, human genomic data and other biomedical data suffer from sampling biases in terms of people's ethnicity, as most of the data come from people of European ancestry. Smaller sample sizes for other population groups can cause suboptimal results in ML-based prediction models for those populations. Suboptimal predictions in precision medicine for some particular group can cause serious consequences limiting the model's applicability in real-world problems. As data collection for those populations is time-consuming and costly, we suggest deep learning-based models for in-silico data enhancement. Existing Generative Adversarial Network (GAN) models for genomic data like Population scale Genomic conditional-GAN (PG-cGAN) can generate realistic genomic data while trained on fairly unbiased data but fails while trained on biased data and encounters severe mode collapse. Our proposed model, Offspring GAN, can resolve the mode collapse issue even when trained in strongly biased genomic datasets. Our results demonstrate the ability of Offspring GAN to generate realistic and diverse label-aware data, which can augment limited real data to alleviate biases and disparities in genomic data. We also propose a privacy-preserving protocol using Offspring GAN to protect the privacy of genomic data.