Solving the protein sequence metric problem

Solving the protein sequence metric problem
复制标题

DOI:
10.1073/pnas.0408677102
复制
发表时间:
2005-05-03
影响因子:
11.1
通讯作者:
Drüke, T
Drüke, T
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Atchley, WR;Zhao, JP;Drüke, T

文献摘要

被引文献

相似文献

生物序列是由字母组成的长串,而不是由数值组成的数组。缺乏一个自然的基本指标比较这样的字母数据显着抑制复杂的序列统计分析,建模蛋白质的结构和功能方面,以及相关的问题。在这里,我们使用多变量统计分析近500个氨基酸属性,以产生一个小的高度可解释的氨基酸变异性的数字模式。这些高维属性数据总结了五个多维模式的属性协变,反映极性,二级结构,分子体积,密码子多样性和静电荷。每个氨基酸的数值评分然后转换氨基酸序列用于统计分析。转换后的数据和氨基酸取代矩阵之间的关系显示极性和密码子多样性得分显着的协会。转换后的字母数据用于方差分析和判别分析,以研究DNA结合的基本螺旋-环-螺旋蛋白。转换后的分数提供了一个通用的解决方案,用于分析各种各样的序列分析问题。
Biological sequences are composed of long strings of alphabetic letters rather than arrays of numerical values. Lack of a natural underlying metric for comparing such alphabetic data significantly inhibits sophisticated statistical analyses of sequences, modeling structural and functional aspects of proteins, and related problems. Herein, we use multivariate statistical analyses on almost 500 amino acid attributes to produce a small set of highly interpretable numeric patterns of amino acid variability. These high-dimensional attribute data are summarized by five multidimensional patterns of attribute covariation that reflect polarity, secondary structure, molecular volume, codon diversity, and electrostatic charge. Numerical scores for each amino acid then transform amino acid sequences for statistical analyses. Relationships between transformed data and amino acid substitution matrices show significant associations for polarity and codon diversity scores. Transformed alphabetic data are used in analysis of variance and discriminant analysis to study DNA binding in the basic helix-loop-helix proteins. The transformed scores offer a general solution for analyzing a wide variety of sequence analysis problems.