STATISTICS OF LOCAL COMPLEXITY IN AMINO-ACID-SEQUENCES AND SEQUENCE DATABASES

STATISTICS OF LOCAL COMPLEXITY IN AMINO-ACID-SEQUENCES AND SEQUENCE DATABASES
复制标题

DOI:
10.1016/0097-8485(93)85006-x
复制
发表时间:
1993-06-01
期刊:
COMPUTERS & CHEMISTRY
影响因子:
--
通讯作者:
FEDERHEN, S
FEDERHEN, S
中科院分区:
其他
文献类型:
--
作者:
WOOTTON, JC;FEDERHEN, S

文献摘要

被引文献

相似文献

蛋白质序列令人惊讶地包含许多低组成复杂性的局部区域。这些包括不同类型的残留簇,其中一些包含均聚物,短期重复或几种残留类型的植物镶嵌物。此处介绍了几种局部复杂性和概率的形式定义,并通过在氨基酸序列和序列数据库中定位此类区域的算法中的效用进行了比较。定义是:--( 1)从枚举派生的先验派生的治疗方法类似于统计力学的处理,(2)对数类似于信息熵的复杂性的对数似然定义,(3)观察到的组成的多项式概率,(4)近似值类似于CHI2统计量,(5)变化系数的修改。这些度量以及一种基于不同偏移序列的自我对准序列的相似性得分的方法被证明在蛋白质序列中与低复杂性区域的近似定位相似,但在应用时,它们在应用时会产生明显不同的结果。在最佳分割算法中。这些比较基于算法(seg)中强大优化启发式方法的选择,旨在将氨基酸序列完全自动分为对比复杂性的子序列。从SwissProt数据库中对丰富的低复杂性段进行了分区之后,其余的高复杂序列集通过一阶随机模型充分近似。
Protein sequences contain surprisingly many local regions of low compositional complexity. These include different types of residue clusters, some of which contain homopolymers, short period repeats or aperiodic mosaics of a few residue types. Several different formal definitions of local complexity and probability are presented here and are compared for their utility in algorithms for localization of such regions in amino acid sequences and sequence databases. The definitions are:-(1) those derived from enumeration a priori by a treatment analogous to statistical mechanics, (2) a log likelihood definition of complexity analogous to informational entropy, (3) multinomial probabilities of observed compositions, (4) an approximation resembling the chi2 statistic and (5) a modification of the coefficient of divergence. These measures, together with a method based on similarity scores of self-aligned sequences at different offsets, are shown to be broadly similar for first-pass, approximate localization of low-complexity regions in protein sequences, but they give significantly different results when applied in optimal segmentation algorithms. These comparisons underpin the choice of robust optimization heuristics in an algorithm, SEG, designed to segment amino acid sequences fully automatically into subsequences of contrasting complexity. After the abundant low-complexity segments have been partitioned from the Swissprot database, the remaining high-complexity sequence set is adequately approximated by a first-order random model.