Kmacs: the k-mismatch average common substring approach to alignment-free sequence comparison.

Kmacs: the k-mismatch average common substring approach to alignment-free sequence comparison.
复制标题

DOI:
10.1093/bioinformatics/btu331
复制
发表时间:
2014-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Morgenstern B
Morgenstern B
中科院分区:
其他
文献类型:
--
作者:
Leimeister CA;Morgenstern B

文献摘要

参考文献

被引文献

相似文献

动机:如果要分析大型数据集,则用于序列分析的基于比对的方法具有各种限制。因此,近年来,无干扰方法变得流行。最著名的无重复方法之一是平均公共子串方法,该方法基于序列之间最长公共词的平均长度来定义序列上的距离度量。在这里,我们通过考虑具有k个不匹配的最长公共子串来推广这种方法。我们提出了一个贪婪的启发式近似这样的k-不匹配的子串的长度,我们描述kmacs,这一想法的基础上广义增强后缀数组的有效实现。结果:为了评估我们的方法的性能,我们将其应用于使用大量DNA和蛋白质序列集的同源性重建。在大多数情况下,用kmacs计算的系统发育树比用基于精确单词匹配的无重复方法生成的树更准确。特别是在蛋白质序列上,我们的方法似乎是上级。在模拟的蛋白质家族中,kmacs的表现甚至超过了使用多重比对和最大似然法进行同源性重建的经典方法。可用性和实现:kmacs是用C++实现的,源代码可以在http://kmacs.gobics.de/上免费获得。联系方式:chris. stud.uni-goettingen.de补充信息:补充数据可以在生物信息学在线上获得。
Motivation: Alignment-based methods for sequence analysis have various limitations if large datasets are to be analysed. Therefore, alignment-free approaches have become popular in recent years. One of the best known alignment-free methods is the average common substring approach that defines a distance measure on sequences based on the average length of longest common words between them. Herein, we generalize this approach by considering longest common substrings with k mismatches. We present a greedy heuristic to approximate the length of such k-mismatch substrings, and we describe kmacs, an efficient implementation of this idea based on generalized enhanced suffix arrays. Results: To evaluate the performance of our approach, we applied it to phylogeny reconstruction using a large number of DNA and protein sequence sets. In most cases, phylogenetic trees calculated with kmacs were more accurate than trees produced with established alignment-free methods that are based on exact word matches. Especially on protein sequences, our method seems to be superior. On simulated protein families, kmacs even outperformed a classical approach to phylogeny reconstruction using multiple alignment and maximum likelihood. Availability and implementation: kmacs is implemented in C++, and the source code is freely available at http://kmacs.gobics.de/ Contact: chris.leimeister@stud.uni-goettingen.de Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1186/1471-2105-14-248
发表时间: 2013-08-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Hauser M;Mayer CE;Söding J
通讯作者: Söding J
DOI: 10.1093/bioinformatics/14.2.157
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stoye, J;Evers, D;Meyer, F
通讯作者: Meyer, F
DOI: 10.1038/ismej.2009.150
发表时间: 2010-06-01
期刊: ISME JOURNAL
影响因子: 11
作者:
Newton, Ryan J.;Griffin, Laura E.;Moran, Mary Ann
通讯作者: Moran, Mary Ann
DOI: 10.1093/bioinformatics/btu177
发表时间: 2014-07-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Leimeister CA;Boden M;Horwege S;Lindner S;Morgenstern B
通讯作者: Morgenstern B
DOI: 10.1186/1471-2105-6-123
发表时间: 2005-05-23
期刊: BMC bioinformatics
影响因子: 3
作者:
Haubold B;Pierstorff N;Möller F;Wiehe T
通讯作者: Wiehe T