Simplified amino acid alphabets for protein fold recognition and implications for folding

Simplified amino acid alphabets for protein fold recognition and implications for folding
复制标题

DOI:
10.1093/protein/13.3.149
复制
发表时间:
2000-03-01
期刊:
PROTEIN ENGINEERING
影响因子:
--
通讯作者:
Levy, RM
Levy, RM
中科院分区:
其他
文献类型:
--
作者:
Murphy, LR;Wallqvist, A;Levy, RM

文献摘要

被引文献

相似文献

蛋白质设计实验表明,使用特定的氨基酸子集可以产生可折叠的蛋白质。这就提出了一个问题,即是否存在一个最小的氨基酸字母表,可以用来折叠所有的蛋白质。在这项工作中,我们做一个类比之间的序列模式,产生可折叠的序列和那些使得有可能通过比对序列来检测结构同源物,并使用它来建议可能的大小,这样一个减少字母表。我们估计,包含10-12个字母的简化字母表可用于设计大量蛋白质家族的可折叠序列。这种估计是基于这样的观察,即当将氨基酸字母表从20个字母适当减少到10个字母时,在聚类的蛋白质序列数据库中挑选出结构同源物所需的信息几乎没有损失,但是当进一步减少字母表时,这种信息迅速退化。
Protein design experiments have shown that the use of specific subsets of amino acids can produce foldable proteins. This prompts the question of whether there is a minimal amino acid alphabet which could be used to fold all proteins. In this work we make an analogy between sequence patterns which produce foldable sequences and those which make it possible to detect structural homologs by aligning sequences, and use it to suggest the possible size of such a reduced alphabet. We estimate that reduced alphabets containing 10-12 letters can be used to design foldable sequences for a large number of protein families. This estimate is based on the observation that there is little loss of the information necessary to pick out structural homologs in a clustered protein sequence database when a suitable reduction of the amino acid alphabet from 20 to 10 letters is made, but that this information is rapidly degraded when further reductions in the alphabet are made.