SEK: sparsity exploiting k-mer-based estimation of bacterial community composition

SEK: sparsity exploiting k-mer-based estimation of bacterial community composition
复制标题

DOI:
10.1093/bioinformatics/btu320
复制
发表时间:
2014-09-01
期刊:
影响因子:
5.8
通讯作者:
Corander, Jukka
Corander, Jukka
中科院分区:
生物学3区
文献类型:
--
作者:
Chatterjee, Saikat;Koslicki, David;Corander, Jukka

文献摘要

被引文献

相似文献

动机:从高通量测序样品中估计细菌群落组成是宏基因组学应用中的一项重要任务。由于样品序列数据通常包含可变长度的读段和不同水平的生物和技术噪声,因此此类数据的准确统计分析具有挑战性。目前流行的估计方法通常是耗时的,在桌面computing environment.Results:使用稀疏执行方法从一般的稀疏信号处理领域(如压缩感知),我们推导出一个解决方案的社区组成估计问题的同时分配的所有样本读取到一个预处理的参考数据库。基于核密度估计技术的一般统计模型的分配任务,并使用凸优化工具获得模型的解决方案。此外,我们设计了一个贪婪算法的快速解决方案。我们的方法提供了一个合理的快速社区组成估计方法,这是更强大的输入数据的变化比最近推出的相关方法。
Motivation: Estimation of bacterial community composition from a high-throughput sequenced sample is an important task in metagenomics applications. As the sample sequence data typically harbors reads of variable lengths and different levels of biological and technical noise, accurate statistical analysis of such data is challenging. Currently popular estimation methods are typically time-consuming in a desktop computing environment.Results: Using sparsity enforcing methods from the general sparse signal processing field (such as compressed sensing), we derive a solution to the community composition estimation problem by a simultaneous assignment of all sample reads to a pre-processed reference database. A general statistical model based on kernel density estimation techniques is introduced for the assignment task, and the model solution is obtained using convex optimization tools. Further, we design a greedy algorithm solution for a fast solution. Our approach offers a reasonably fast community composition estimation method, which is shown to be more robust to input data variation than a recently introduced related method.