Phylogenetic double placement of mixed samples

Phylogenetic double placement of mixed samples
复制标题

DOI:
10.1093/bioinformatics/btaa489
复制
发表时间:
2020-07-01
期刊:
影响因子:
5.8
通讯作者:
Mirarab,Siavash
Mirarab,Siavash
中科院分区:
生物学3区
文献类型:
--
作者:
Balaban,Metin;Mirarab,Siavash

文献摘要

被引文献

相似文献

动机考虑一个简单的计算问题。输入是(i)从组合两种生物体的样品产生的混合读数的集合和(ii)已知来源的若干参考基因组的单独读数的集合。目标是找到构成混合样品的两种生物体。当成分是从参考集缺席,我们试图将它们相对于参考物种的基础树的遗传位置。这个简单而又基本的问题(我们称之为系统发育双重定位)在文献中很少受到关注。作为基因组略读(低通测序的基因组在低覆盖率,排除组装)变得越来越普遍,这个问题发现广泛的应用领域,如生物多样性研究,食品生产和原产地,和进化reconstruction.ResultsWe引入一个模型,涉及混合样品和参考物种之间的距离成分和参考物种之间的距离。我们的模型是基于Jaccard指数计算的每个样本之间表示为k-mer集。该模型,建立在几个假设和近似值,使我们能够正式的系统发育双布局问题作为一个非凸优化问题,分解混合物的距离,同时进行系统发育的位置。使用各种技术,我们能够解决这个优化问题的数值。我们测试所得到的方法,称为混合样本分析工具(MISA),在一组不同的模拟和生物数据集。尽管使用了所有的假设,该方法在实践中表现得非常好。可用性和实施该软件和数据可在https://github.com/balabanmetin/misa和https://github.com/balabanmetin/misa-data.Supplementary信息补充数据可在生物信息学在线。
MotivationConsider a simple computational problem. The inputs are (i) the set of mixed reads generated from a sample that combines two organisms and (ii) separate sets of reads for several reference genomes of known origins. The goal is to find the two organisms that constitute the mixed sample. When constituents are absent from the reference set, we seek to phylogenetically position them with respect to the underlying tree of the reference species. This simple yet fundamental problem (which we call phylogenetic double-placement) has enjoyed surprisingly little attention in the literature. As genome skimming (low-pass sequencing of genomes at low coverage, precluding assembly) becomes more prevalent, this problem finds wide-ranging applications in areas as varied as biodiversity research, food production and provenance, and evolutionary reconstruction.ResultsWe introduce a model that relates distances between a mixed sample and reference species to the distances between constituents and reference species. Our model is based on Jaccard indices computed between each sample represented as k-mer sets. The model, built on several assumptions and approximations, allows us to formalize the phylogenetic double-placement problem as a non-convex optimization problem that decomposes mixture distances and performs phylogenetic placement simultaneously. Using a variety of techniques, we are able to solve this optimization problem numerically. We test the resulting method, called MIxed Sample Analysis tool (MISA), on a varied set of simulated and biological datasets. Despite all the assumptions used, the method performs remarkably well in practice.Availability and implementationThe software and data are available at https://github.com/balabanmetin/misa and https://github.com/balabanmetin/misa-data.Supplementary informationSupplementary data are available atBioinformaticsonline.