pong: fast analysis and visualization of latent clusters in population genetic data

pong: fast analysis and visualization of latent clusters in population genetic data
复制标题

DOI:
10.1093/bioinformatics/btw327
复制
发表时间:
2016-09-15
期刊:
影响因子:
5.8
通讯作者:
Ramachandran, Sohini
Ramachandran, Sohini
中科院分区:
生物学3区
文献类型:
--
作者:
Behr, Aaron A.;Liu, Katherine Z.;Ramachandran, Sohini

文献摘要

被引文献

相似文献

动机:群体遗传学中的一系列方法使用多位点基因型数据来分配个体在潜在聚类中的成员资格。这些方法属于广泛的一类混合成员模型,如用于分析文本语料库的潜在狄利克雷分配。当重复应用于相同的输入时,来自混合成员模型的推断可以产生不同的输出矩阵,并且潜在聚类的数量是在分析管道中经常变化的参数。出于这些原因,量化,可视化和注释混合成员模型的输出是瓶颈,跨多个学科的研究人员从生态学到文本数据mining.Results:我们介绍了乒乓,网络图形化的方法,用于分析和可视化的成员在潜在的集群与本地交互式D3.js可视化。pong利用高效的算法来解决分配问题,与处理来自混合成员模型的输出的其他方法相比,大大减少了运行时间,同时提高了准确性。我们将pong应用于来自1000个基因组计划中2426个无关个体的225 705个未链接的全基因组单核苷酸变体,并识别出全球人类人口结构中以前被忽视的方面。我们表明,乒乓超越当前的解决方案在运行时超过一个数量级,同时提供了一个可定制的和交互式的可视化的人口结构,比目前的工具产生的更准确。
Motivation: A series of methods in population genetics use multilocus genotype data to assign individuals membership in latent clusters. These methods belong to a broad class of mixed-membership models, such as latent Dirichlet allocation used to analyze text corpora. Inference from mixed- membership models can produce different output matrices when repeatedly applied to the same inputs, and the number of latent clusters is a parameter that is often varied in the analysis pipeline. For these reasons, quantifying, visualizing, and annotating the output from mixed-membership models are bottlenecks for investigators across multiple disciplines from ecology to text data mining.Results: We introduce pong, a network-graphical approach for analyzing and visualizing membership in latent clusters with a native interactive D3.js visualization. pong leverages efficient algorithms for solving the Assignment Problem to dramatically reduce runtime while increasing accuracy compared with other methods that process output from mixed-membership models. We apply pong to 225 705 unlinked genome-wide single-nucleotide variants from 2426 unrelated individuals in the 1000 Genomes Project, and identify previously overlooked aspects of global human population structure. We show that pong outpaces current solutions by more than an order of magnitude in runtime while providing a customizable and interactive visualization of population structure that is more accurate than those produced by current tools.