Tractable and Expressive Generative Models of Genetic Variation Data.

Tractable and Expressive Generative Models of Genetic Variation Data.
复制标题

遗传变异数据的易处理且富有表现力的生成模型。

DOI:
10.1101/2023.05.16.541036
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
VandenBroeck,Guy
VandenBroeck,Guy
中科院分区:
--
文献类型:
--
作者:
Dang,Meihua;Liu,Anji;Wei,Xinzhu;Sankararaman,Sriram;VandenBroeck,Guy

文献摘要

相似文献

群体遗传学研究通常依赖于由遗传数据的生成模型模拟的人工基因组(AG)。近年来,基于隐马尔可夫模型、深度生成对抗网络、受限玻尔兹曼机和变分自编码器的无监督学习模型由于其生成与经验数据非常相似的AG的能力而受到欢迎。然而,这些模型在表达性和易处理性之间进行了权衡。在这里,我们建议使用隐藏的Chow-Liu树(HCLTs)及其表示为概率电路(PC)作为这种权衡的解决方案。我们首先学习HCLT结构,该结构捕获训练数据集中SNP之间的长程依赖关系。然后,我们将HCLT转换为等价的PC作为一种手段,支持易处理和有效的概率推理。这些PC中的参数推断与期望最大化算法使用的训练数据。与用于生成AG的其他模型相比,HCLT在跨基因组和从连续基因组区域选择的SNP的测试基因组上获得最大的对数似然。此外,HCLT产生的AG更准确地类似于源数据集的等位基因频率,连锁不平衡,成对单倍型距离和人口结构的模式。这项工作不仅提出了一个新的和强大的AG模拟器,但也体现了潜在的PC在群体遗传学。
Population genetic studies often rely on artificial genomes (AGs) simulated by generative models of genetic data. In recent years, unsupervised learning models, based on hidden Markov models, deep generative adversarial networks, restricted Boltzmann machines, and variational autoencoders, have gained popularity due to their ability to generate AGs closely resembling empirical data. These models, however, present a tradeoff between expressivity and tractability. Here, we propose to use hidden Chow-Liu trees (HCLTs) and their representation as probabilistic circuits (PCs) as a solution to this tradeoff. We first learn an HCLT structure that captures the long-range dependencies among SNPs in the training data set. We then convert the HCLT to its equivalent PC as a means of supporting tractable and efficient probabilistic inference. The parameters in these PCs are inferred with an expectation-maximization algorithm using the training data. Compared to other models for generating AGs, HCLT obtains the largest log-likelihood on test genomes across SNPs chosen across the genome and from a contiguous genomic region. Moreover, the AGs generated by HCLT more accurately resemble the source data set in their patterns of allele frequencies, linkage disequilibrium, pairwise haplotype distances, and population structure. This work not only presents a new and robust AG simulator but also manifests the potential of PCs in population genetics.