Fast alignment-free sequence comparison using spaced-word frequencies.

Fast alignment-free sequence comparison using spaced-word frequencies.
复制标题

DOI:
10.1093/bioinformatics/btu177
复制
发表时间:
2014-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Morgenstern B
Morgenstern B
中科院分区:
其他
文献类型:
--
作者:
Leimeister CA;Boden M;Horwege S;Lindner S;Morgenstern B

文献摘要

被引文献

相似文献

动机:用于序列比较的免比对方法越来越多地用于基因组分析和系统发育重建;它们规避了传统的基于对齐的方法的各种困难。特别是,免比对方法比成对或多重比对快得多。然而,它们的准确性不如基于序列比对的方法。大多数免对齐方法通过比较序列的单词组成来工作。这些方法的一个众所周知的问题是相邻单词匹配远非独立。结果:为了减少相邻单词匹配之间的统计依赖性,我们建议使用由“匹配”和“不关心”位置模式定义的“间隔单词”进行无对齐序列比较。我们描述了使用递归散列和位操作的这种方法的快速实现,并且我们表明可以通过使用多个模式而不是单个模式来实现进一步的改进。为了评估我们的方法,我们使用间隔词频率作为快速系统发育重建的基础。使用真实世界和模拟序列数据,我们证明我们的多模式方法比依赖连续词的方法产生更好的系统发育。可用性和实施​​:我们的计划可在 http://spaced.gobics.de/ 上免费获取。联系方式:chris.leimeister@stud.uni-goettingen.de 补充信息:补充数据可在生物信息学在线获取。
Motivation: Alignment-free methods for sequence comparison are increasingly used for genome analysis and phylogeny reconstruction; they circumvent various difficulties of traditional alignment-based approaches. In particular, alignment-free methods are much faster than pairwise or multiple alignments. They are, however, less accurate than methods based on sequence alignment. Most alignment-free approaches work by comparing the word composition of sequences. A well-known problem with these methods is that neighbouring word matches are far from independent. Results: To reduce the statistical dependency between adjacent word matches, we propose to use ‘spaced words’, defined by patterns of ‘match’ and ‘don’t care’ positions, for alignment-free sequence comparison. We describe a fast implementation of this approach using recursive hashing and bit operations, and we show that further improvements can be achieved by using multiple patterns instead of single patterns. To evaluate our approach, we use spaced-word frequencies as a basis for fast phylogeny reconstruction. Using real-world and simulated sequence data, we demonstrate that our multiple-pattern approach produces better phylogenies than approaches relying on contiguous words. Availability and implementation: Our program is freely available at http://spaced.gobics.de/. Contact: chris.leimeister@stud.uni-goettingen.de Supplementary information: Supplementary data are available at Bioinformatics online.