Sequence alignment:: an approximation law for the Z-value with applications to databank scanning

Sequence alignment:: an approximation law for the Z-value with applications to databank scanning
复制标题

DOI:
10.1016/s0097-8485(01)00074-2
复制
发表时间:
2001-07-01
期刊:
COMPUTERS & CHEMISTRY
影响因子:
--
通讯作者:
Comet, JP
Comet, JP
中科院分区:
其他
文献类型:
--
作者:
Bacro, JN;Comet, JP

文献摘要

被引文献

相似文献

Z值是通过使用蒙特卡罗程序来估计史密斯和沃特曼动态规划比对分数(H分数)的统计意义的尝试。本文给出了由Watman和Vingron(Stat.)提出的泊松聚集启发式推导出的Z值律的近似解。SCI。9(1994)367)在独立且同分布的序列比较的情况下。对于无间隙比对得分,我们的近似为Gumbel类型,但参数与序列无关。这一结果澄清了彗星等人提到的相关实验结果。(计算机。化学。23(1999)317)。使用‘准实’序列(即与真实序列长度和氨基酸组成相同的随机洗牌序列),我们研究了我们的近似结果的相关性。针对蒙特卡罗方法对Gumbel衰变参数估计产生偏差的问题,提出了一种修正方法。我们考虑了对真实序列的应用,并展示了我们的结果如何用于检测真实序列之间的潜在生物关系。(C)2001爱思唯尔科学有限公司。保留所有权利。
The Z-value is an attempt to estimate the statistical significance of a Smith and Waterman dynamic programming alignment score (H-score) through the use of a Monte-Carlo procedure. In this paper, we give an approximation for the Z-value law deduced from the Poisson clumping heuristic developed by Waterman and Vingron (Stat. Sci. 9 (1994) 367) in the case of independent and identically distributed sequences comparison. As for non-gapped alignment scores, our approximation is of Gumbel type but with parameters that are sequence independent. This result makes clear the related experimental results mentioned by Comet et al. (Comput. Chem. 23 (1999) 317). Using 'quasi-real' sequences (i.e. randomly shuffled sequences of the same length and amino acid composition as the real ones) we investigate the relevance of our approximation result. Since the Monte-Carlo approach we use generates a bias for the Gumbel decay parameter estimation, a correction procedure is proposed. Applications to real sequences are considered and we show how our results can be used to detect the potential biological relationships between real sequences. (C) 2001 Elsevier Science Ltd. All rights reserved.