METHODS FOR ASSESSING THE STATISTICAL SIGNIFICANCE OF MOLECULAR SEQUENCE FEATURES BY USING GENERAL SCORING SCHEMES

METHODS FOR ASSESSING THE STATISTICAL SIGNIFICANCE OF MOLECULAR SEQUENCE FEATURES BY USING GENERAL SCORING SCHEMES
复制标题

DOI:
10.1073/pnas.87.6.2264
复制
发表时间:
1990-03-01
影响因子:
11.1
通讯作者:
ALTSCHUL, SF
ALTSCHUL, SF
中科院分区:
综合性期刊1区
文献类型:
--
作者:
KARLIN, S;ALTSCHUL, SF

文献摘要

被引文献

相似文献

核酸或蛋白质序列中的异常模式或两个或多个序列共享的强相似性区域可能具有生物学意义。因此,人们希望知道这种模式是否仅仅是偶然出现的。为了鉴定感兴趣的序列模式,当比较几个序列时,可以将适当的评分值分配给单个序列的各个残基或残基组。对于单个序列,这样的分数可以反映生物物理特性,例如电荷、体积、疏水性或二级结构潜力;对于多个序列,它们可以反映以多种方式测量的核苷酸或氨基酸相似性。使用一个适当的随机模型,我们提出了一个理论,提供精确的数值公式,用于评估任何地区的统计显著性与高的总得分。第二类结果描述了高分段的组成。在某些情况下,这些允许选择的评分系统是“最佳的”区分生物相关的模式。例子给出了各种蛋白质序列的理论的应用,突出不寻常的生物学特征的片段。这些包括转录因子和原癌基因产物中独特的电荷区域,各种受体和转运蛋白中明显的疏水片段,以及涉及最近表征的囊性纤维化基因的统计学显著的亚对齐。
An unusual pattern in a nucleic acid or protein sequence or a region of strong similarity shared by two or more sequences may have biological significance. It is therefore desirable to known whether such a pattern can have arisen simply by chance. To identify interesting sequence patterns, appropriate scoring values can be assigned to the individual residues of a single sequence or to sets of residues when several sequences are compared. For single sequences, such scores can reflect biophysical properties such as charge, volume, hydrophobicity, or secondary structure potential; for multiple sequences, they can reflect nucleotide or amino acid similarity measured in a wide variety of ways. Using an appropriate random model, we present a theory that provides precise numerical formulas for assessing the statistical significance of any region with high aggregate score. A second class of results describes the composition of high-scoring segments. In certain contexts, these permit the choice of scoring systems which are "optimal" for distinguishing biologically relevant patterns. Examples are given of applications of the theory to a variety of protein sequences, highlighting segments with unusual biological features. These include distinctive charge regions in transcription factors and protooncogene products, pronounced hydrophobic segments in various receptor and transport proteins, and statistically significant subalignments involving the recently characterized cystic fibrosis gene.