A NONLINEAR MEASURE OF SUBALIGNMENT SIMILARITY AND ITS SIGNIFICANCE LEVELS

A NONLINEAR MEASURE OF SUBALIGNMENT SIMILARITY AND ITS SIGNIFICANCE LEVELS
复制标题

DOI:
10.1007/bf02462327
复制
发表时间:
1986-01-01
影响因子:
3.5
通讯作者:
ERICKSON, BW
ERICKSON, BW
中科院分区:
数学4区
文献类型:
--
作者:
ALTSCHUL, SF;ERICKSON, BW

文献摘要

被引文献

相似文献

提出了一种新的亚列相似度度量方法。具体来说,相似性s(l, c)被定义为在长度为l的子序列中找到c或更少不匹配的概率的以p为底的对数,其中p是匹配的概率。以前的算法不能使用这种度量来找到局部最优的子排列,因为,与Needleman-Wunsch和sellers相似度不同,这种度量是非线性的。描述了一种新的模式识别算法,用于寻找两个核苷酸序列的所有局部最优亚对。DD算法可以使用s(l, c)或任何其他合理的相似度函数来评估子序列的相对兴趣。DD算法只搜索对角线图,缺少插入和删除。这种搜索策略大大减少了计算时间,并且不需要任意选择间隙代价。生成的DD图的路径通常会注意到插入和删除的可能位置。推导了一个启发式公式,用于在两个对齐序列的长度上下文中估计s(l, c)的显著性水平。DD算法已被用于发现人类和小鼠白细胞介素2核苷酸序列之间有趣的亚序列。
A new measure of subalignment similarity is introduced. Specifically, similarity s(l, c) is defined as the logarithm to the base p of the probability of finding c or fewer mismatches in a subalignment of length l, where p is the probability of a match. Previous algorithms can not use this measure to find locally optimal subalignments because, unlike Needleman-Wunsch and sellers similarities, this measure is nonlinear. A new pattern recognition algorithm is described for finding all locally optimal subalignments of two nucleotide sequences. The DD algorithm can use s(l, c) or any other reasonable similarity function to assess the relative interest of subalignments. The DD lgorithm searches only the diagonal graph, which lacks insertions and deletions. This search strategy greatly decreases the computation time and does not require an arbitrary choice of gap cost. The paths of the resulting DD graph usually draw attention to likely locations for insertions and deletions. A heuristic formula is derived for estimating significance levels for s(l, c) in the context of the lengths of the two aligned sequences. The DD algorithm has been used to find interesting subalignments between the nucleotide sequences for human and murine interleukin 2.