Population-scale Genomic Data Augmentation Based on Conditional Generative Adversarial Networks

Population-scale Genomic Data Augmentation Based on Conditional Generative Adversarial Networks
复制标题

DOI:
10.1145/3388440.3412475
复制
发表时间:
2020-09
期刊:
Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics
影响因子:
--
通讯作者:
Junjie Chen;M. Mowlaei;Xinghua Shi
Junjie Chen;M. Mowlaei;Xinghua Shi
中科院分区:
其他
文献类型:
--
作者:
Junjie Chen;M. Mowlaei;Xinghua Shi

文献摘要

被引文献

相似文献

尽管下一代测序技术已经使得快速生成大量序列成为可能,但由于疾病罕见、测试负担能力以及对隐私和安全的担忧等多种因素,当前的基因组数据仍然存在数据量小、不平衡和偏差等问题。为了解决基因组数据的这些局限性,我们开发了基于条件生成对抗网络(PG-cGAN)的群体规模基因组数据增强,通过转换数据中已有的样本而不是收集新样本来增强基因组数据的数量和多样性。 PG-CGAN 中的生成器和鉴别器都堆叠有卷积层,以捕获底层的群体结构。我们在人类白细胞抗原 (HLA) 区域中增强基因型的结果表明,PC-cGAN 可以生成具有相似群体结构、变异频率分布和 LD 模式的新基因型。由于 PC-cGAN 的输入是原始基因组数据,没有对先验知识的假设,因此它可以扩展以丰富许多其他类型的生物医学数据及其他数据。
Although next generation sequencing technologies have made it possible to quickly generate a large collection of sequences, current genomic data still suffer from small data sizes, imbalances, and biases due to various factors including disease rareness, test affordability, and concerns about privacy and security. In order to address these limitations of genomic data, we develop a Population-scale Genomic Data Augmentation based on Conditional Generative Adversarial Networks (PG-cGAN) to enhance the amount and diversity of genomic data by transforming samples already in the data rather than collecting new samples. Both the generator and discriminator in the PG-CGAN are stacked with convolutional layers to capture the underlying population structure. Our results for augmenting genotypes in human leukocyte antigen (HLA) regions showed that PC-cGAN can generate new genotypes with similar population structure, variant frequency distributions and LD patterns. Since the input for PC-cGAN is the original genomic data without assumptions about prior knowledge, it can be extended to enrich many other types of biomedical data and beyond.