Crystal structure of a novel non-Pfam protein AF1514 from Archeoglobus fulgidus DSM 4304 solved by S-SAD using a Cr X-ray source.

Crystal structure of a novel non-Pfam protein AF1514 from Archeoglobus fulgidus DSM 4304 solved by S-SAD using a Cr X-ray source.
复制标题

使用 Cr X 射线源通过 S-SAD 解析来自 Archeoglobus fulgidus DSM 4304 的新型非 Pfam 蛋白 AF1514 的晶体结构。

DOI:
10.1002/prot.22025
复制
发表时间:
2008
期刊:
影响因子:
2.9
通讯作者:
Liu,Zhi-Jie
Liu,Zhi-Jie
中科院分区:
生物学4区
文献类型:
--
作者:
Li,Yang;Bahti,Pazilat;Shaw,Neil;Song,Gaojie;Chen,Shunmei;Zhang,Xuejun;Zhang,Min;Cheng,Chongyun;Yin,Jie;Zhu,Jin-Yi;Zhang,Hua;Che,Dongsheng;Xu,Hao;Abbas,Abdulla;Wang,Bi-Cheng;Liu,Zhi-Jie

文献摘要

相似文献

最近开发了许多计算工具来根据蛋白质的氨基酸序列准确地预测蛋白质的结构。1-3通常,如果查询蛋白质序列与其3D结构已被确定的蛋白质序列具有至少30%的序列同一性,则该查询序列的结构可以基于模板结构来建模,例如使用建模器软件。然而,3、4计算软件不能保证对PDB中序列相似性较低的新蛋白质的准确预测。因此,X射线结晶学和核磁共振等实验方法仍然是蛋白质结构研究的主要手段。5、6目前大多数结构基因组学中心的靶点选择策略7-9主要集中在人工调控蛋白家族(PFAM)的代表,10-12,即所选择的蛋白质序列与家族内的其他成员至少有一个保守结构域。这样,求解的代表性结构可以作为结构模板,使用计算工具来预测同一家族中剩余蛋白质序列的结构。结果表明,这种“Pfam”靶选择策略不仅增加了新结构的数量,而且增加了新折叠的数量。13然而,在靶点选择中过分强调Pfam而忽略非Pfam序列(即不共享Pfam中的任何保守结构域),可能会导致Pfam和非Pfam结构在PDB中的偏向分布,并可能减缓新结构和折叠的生长速度。我们对150个微生物基因组的分析表明,大多数基因组的非Pfam序列占所有开放阅读框架(ORF)的25%-30%,有些可达到60%(未发表的数据)。非私营部门金融机构占所有非私营部门金融机构的比例很高,这提醒我们,在制定目标选择战略时,不应忽视非私营部门金融机构。另一方面,每个基因组的这些非Pfam序列要么是平行的非Pfam(其中序列在同一有机体中具有同源伙伴),要么是同源的非Pfam(其中序列在紧密地
Many computational tools have been developed recently to accurately predict the structure of a protein from its amino acid sequence. 1–3 In general, if a query protein sequence shares at least 30% sequence identity with a protein sequence whose 3D structure has been determined, the structure of this query sequence can be modeled based on the template structure, using MODELLER software for example. 3, 4 Computational software, however, cannot guarantee accurate prediction for those new proteins that share low sequence similarity in PDB. Therefore, experimental methods such as X-ray crystallography and nuclear magnetic resonance (NMR) are still the main approaches for a protein structural study. 5, 6 The current target selection strategy of most structural genomics centers7–9 mainly focuses on the representatives of manually curated protein families (Pfam), 10–12 that is, the selected protein sequence shares at least one conserved domain with other members within a family. In this way, the solved representative structures can be used as structural templates to predict structures of the remaining protein sequences in the same family using computational tools. It has been shown that this ‘‘Pfam’’target selection strategy increases not only the number of novel structures, but also the number of new folds. 13 However, over-emphasis of Pfam and ignoring non-Pfam sequences (ie, not sharing any conserved domain in Pfam) in target selection might lead to biased distribution of Pfam and non-Pfam structures in PDB, and possibly slow the growth rate of new structures and folds. Our analysis on 150 microbial genomes showed that non-Pfam sequences account for 25–30% of all Open Reading Frames (ORFs) for most genomes, and some could reach to 60%(unpublished data). The high percentage of non-Pfams over all ORFs reminds us non-Pfams should not be neglected while devising a target selection strategy. On the other hand, these non-Pfam sequences for each genome are either paralogous non-Pfam (in which sequences have homologous partners within the same organism), or orthologous non-Pfam (in which sequences have orthologous partners in the closely