Finer Metagenomic Reconstruction via Biodiversity Optimization

Finer Metagenomic Reconstruction via Biodiversity Optimization
复制标题

DOI:
10.1101/2020.01.23.916924
复制
发表时间:
2020-01
期刊:
bioRxiv
影响因子:
--
通讯作者:
S. Foucart;D. Koslicki
S. Foucart;D. Koslicki
中科院分区:
其他
文献类型:
--
作者:
S. Foucart;D. Koslicki

文献摘要

相似文献

当从测序的DNA中分析微生物群落时,一项重要的任务是分类特征分析:列举样品中包含的所有生物体或仅仅所有分类群的存在和相对丰度。这项任务可以通过基于压缩传感的方法来解决,这种方法有利于在与观察到的DNA数据一致的生物体中具有最少生物体的群落。尽管取得了成功,但这些简约的方法有时会因忽视生物体的相似性而与生物现实主义相冲突。在这里,我们利用最近开发的生物多样性的概念,同时考虑生物体的相似性,并保留基于压缩传感的方法的优化策略。我们证明,最大限度地减少生物多样性仍然产生稀疏的分类配置文件,我们实验验证现有的压缩传感为基础的方法的优越性。尽管表明,目标函数几乎从来没有凸,往往凹,一般产生NP-困难的问题,我们表现出的方式来表示生物体的相似性,最小化的多样性可以通过一系列的线性规划,保证减少多样性。更好的是,当生物相似性通过k-mer共现(生物信息学中的一个流行概念)量化时,最小化多样性实际上简化为一个线性程序,可以利用多个k-mer大小来提高性能。在概念验证实验中,我们验证了后一种方法在对宏基因组样本进行分类学分析时,无论是在重建精度还是计算性能方面,都可以获得显着的收益。可复制代码可在https://github.com/dkoslicki/MinimizeBiologicalDiversity上获得。
When analyzing communities of microorganisms from their sequenced DNA, an important task is taxonomic profiling: enumerating the presence and relative abundance of all organisms, or merely of all taxa, contained in the sample. This task can be tackled via compressive-sensing-based approaches, which favor communities featuring the fewest organisms among those consistent with the observed DNA data. Despite their successes, these parsimonious approaches sometimes conflict with biological realism by overlooking organism similarities. Here, we leverage a recently developed notion of biological diversity that simultaneously accounts for organism similarities and retains the optimization strategy underlying compressive-sensing-based approaches. We demonstrate that minimizing biological diversity still produces sparse taxonomic profiles and we experimentally validate superiority to existing compressive-sensing-based approaches. Despite showing that the objective function is almost never convex and often concave, generally yielding NP-hard problems, we exhibit ways of representing organism similarities for which minimizing diversity can be performed via a sequence of linear programs guaranteed to decrease diversity. Better yet, when biological similarity is quantified by k-mer co-occurrence (a popular notion in bioinformatics), minimizing diversity actually reduces to one linear program that can utilize multiple k-mer sizes to enhance performance. In proof-of-concept experiments, we verify that the latter procedure can lead to significant gains when taxonomically profiling a metagenomic sample, both in terms of reconstruction accuracy and computational performance. Reproducible code is available at https://github.com/dkoslicki/MinimizeBiologicalDiversity.