THE RAPID GENERATION OF MUTATION DATA MATRICES FROM PROTEIN SEQUENCES

THE RAPID GENERATION OF MUTATION DATA MATRICES FROM PROTEIN SEQUENCES
复制标题

DOI:
10.1093/bioinformatics/8.3.275
复制
发表时间:
1992-06-01
期刊:
COMPUTER APPLICATIONS IN THE BIOSCIENCES
影响因子:
--
通讯作者:
THORNTON, JM
THORNTON, JM
中科院分区:
其他
文献类型:
--
作者:
JONES, DT;TAYLOR, WR;THORNTON, JM

文献摘要

被引文献

相似文献

本文提出了一种从大量蛋白质序列中生成突变数据矩阵的有效方法。通过基于近似肽的序列比较算法,在85%同一性水平下对序列集进行聚类。将最接近的相关序列对进行比对,并将观察到的氨基酸交换记录在矩阵中。原始突变频率矩阵以与Dayhoff等人(1978)描述的方式类似的方式处理,并且因此所得矩阵可以容易地用于当前序列分析应用中,代替标准突变数据矩阵,标准突变数据矩阵已经13年没有更新。该方法是足够快的,以处理整个SWISS-PROT数据库在20小时内在太阳SPARCstation 1,是足够快的,以产生一个矩阵从一个特定的家庭或类的蛋白质在几分钟内。我们的250 PAM突变数据矩阵和Dayhoff等人计算的矩阵之间观察到的差异进行了简要讨论。
An efficient means for generating mutation data matrices from large numbers of protein sequences is presented here. By means of an approximate peptide-based sequence comparison algorithm, the set sequences are clustered at the 85 % identity level. The closest relating pairs of sequences are aligned, and observed amino acid exchanges tallied in a matrix. The raw mutation frequency matrix is processed in a similar way to that described by Dayhoff et al. (1978), and so the resulting matrices may be easily used in current sequence analysis applications, in place of the standard mutation data matrices, which have not been updated for 13 years. The method is fast enough to process the entire SWISS-PROT databank in 20 h on a Sun SPARCstation 1, and is fast enough to generate a matrix from a specific family or class of proteins in minutes. Differences observed between our 250 PAM mutation data matrix and the matrix calculated by Dayhoff et al. are briefly discussed.