Alignment-free genetic sequence comparisons: a review of recent approaches by word analysis

Alignment-free genetic sequence comparisons: a review of recent approaches by word analysis
复制标题

DOI:
10.1093/bib/bbt052
复制
发表时间:
2014-11-01
影响因子:
9.5
通讯作者:
Bastola, Dhundy
Bastola, Dhundy
中科院分区:
生物学2区
文献类型:
--
作者:
Bonham-Carter, Oliver;Steele, Joe;Bastola, Dhundy

文献摘要

被引文献

相似文献

现代测序和基因组组装技术提供了丰富的数据,这些数据很快就需要通过比较分析来发现。序列比对是生物信息学研究中的一项基本任务,可以使用,但有一些注意事项。动态规划的种子技术和方法被证明是无效的这项工作,由于其固有的计算费用时,处理大量的序列数据。由于基因重组、基因改组和其他固有的生物学事件,这些方法容易给出误导性信息。信息论、频率分析和数据压缩的新方法是可用的,并为动态规划提供了强有力的替代方案。这些新的方法往往是首选,因为它们的算法更简单,不受synteny-related problems.In这篇综述中,我们提供了一个详细的讨论的计算工具,这源于无干扰的方法的基础上,从词频的统计分析。我们提供了几个清晰的例子来展示应用程序和解释在几个不同领域的无干扰分析,如基地基地的相关性,特征频率配置文件,组成向量,改进的字符串组成和D-2统计度量。此外,我们提供了详细的讨论和Lempel-Ziv技术从数据压缩分析的例子。
Modern sequencing and genome assembly technologies have provided a wealth of data, which will soon require an analysis by comparison for discovery. Sequence alignment, a fundamental task in bioinformatics research, may be used but with some caveats. Seminal techniques and methods from dynamic programming are proving ineffective for this work owing to their inherent computational expense when processing large amounts of sequence data. These methods are prone to giving misleading information because of genetic recombination, genetic shuffling and other inherent biological events. New approaches from information theory, frequency analysis and data compression are available and provide powerful alternatives to dynamic programming. These new methods are often preferred, as their algorithms are simpler and are not affected by synteny-related problems.In this review, we provide a detailed discussion of computational tools, which stem from alignment-free methods based on statistical analysis from word frequencies. We provide several clear examples to demonstrate applications and the interpretations over several different areas of alignment-free analysis such as base-base correlations, feature frequency profiles, compositional vectors, an improved string composition and the D-2 statistic metric. Additionally, we provide detailed discussion and an example of analysis by Lempel-Ziv techniques from data compression.