An improved string composition method for sequence comparison.

An improved string composition method for sequence comparison.
复制标题

DOI:
10.1186/1471-2105-9-s6-s15
复制
发表时间:
2008-05-28
期刊:
影响因子:
3
通讯作者:
Fang X
Fang X
中科院分区:
生物学4区
文献类型:
--
作者:
Lu G;Zhang S;Fang X

文献摘要

被引文献

相似文献

历史上,两类计算算法(基于比对和无比对)已经应用于序列比较——生物信息学中最基本的问题之一。多序列比对虽然被生物学家广泛使用,但在基础和计算上都有局限性。因此,无比对方法作为估计序列相似性的重要替代方法已被探索。在无比对方法中,利用核苷酸或氨基酸串的频率来表示序列信息的串组成向量(CV)方法在原核生物基因组序列比较中显示出令人满意的结果。然而,现有的基于cv的方法存在一定的统计问题,因此低估了基因序列中进化信息的数量。我们发现现有的基于字符串组合的方法存在两个问题,一个与马尔可夫模型假设有关,另一个与频率归一化方程的分母有关。提出了一种改进的完全组合向量法,在统一独立模型的假设下估计序列信息,有助于序列比较的选择。使用模拟和实验数据集进行的系统发育分析表明,我们的新方法与现有的同类方法相比更具鲁棒性,并且在鲁棒性方面可与基于比对的方法相媲美。针对目前使用的字符串组合方法存在的两个问题,提出了一种新的鲁棒遗传序列进化信息估计方法。此外,我们还讨论了可能没有必要使用相对较长的字符串来构建完整的组合向量(CCV),因为具有可变长度的向量字符串具有重叠性质。我们提出了一种实用的方法来选择构建CCV的最佳字符串长度。
Historically, two categories of computational algorithms (alignment-based and alignment-free) have been applied to sequence comparison–one of the most fundamental issues in bioinformatics. Multiple sequence alignment, although dominantly used by biologists, possesses both fundamental as well as computational limitations. Consequently, alignment-free methods have been explored as important alternatives in estimating sequence similarity. Of the alignment-free methods, the string composition vector (CV) methods, which use the frequencies of nucleotide or amino acid strings to represent sequence information, show promising results in genome sequence comparison of prokaryotes. The existing CV-based methods, however, suffer certain statistical problems, thereby underestimating the amount of evolutionary information in genetic sequences. We show that the existing string composition based methods have two problems, one related to the Markov model assumption and the other associated with the denominator of the frequency normalization equation. We propose an improved complete composition vector method under the assumption of a uniform and independent model to estimate sequence information contributing to selection for sequence comparison. Phylogenetic analyses using both simulated and experimental data sets demonstrate that our new method is more robust compared with existing counterparts and comparable in robustness with alignment-based methods. We observed two problems existing in the currently used string composition methods and proposed a new robust method for the estimation of evolutionary information of genetic sequences. In addition, we discussed that it might not be necessary to use relatively long strings to build a complete composition vector (CCV), due to the overlapping nature of vector strings with a variable length. We suggested a practical approach for the choice of an optimal string length to construct the CCV.