Crystal structure of a novel non-Pfam protein AF1514 from Archeoglobus fulgidus DSM 4304 solved by S-SAD using a Cr X-ray source.
Crystal structure of a novel non-Pfam protein AF1514 from Archeoglobus fulgidus DSM 4304 solved by S-SAD using a Cr X-ray source.
复制标题
使用 Cr X 射线源通过 S-SAD 解析来自 Archeoglobus fulgidus DSM 4304 的新型非 Pfam 蛋白 AF1514 的晶体结构。
DOI:
10.1002/prot.22025
复制
发表时间:
2008
期刊:
影响因子:
2.9
通讯作者:
Liu,Zhi-Jie
中科院分区:
文献类型:
--
作者:
Li,Yang;Bahti,Pazilat;Shaw,Neil;Song,Gaojie;Chen,Shunmei;Zhang,Xuejun;Zhang,Min;Cheng,Chongyun;Yin,Jie;Zhu,Jin-Yi;Zhang,Hua;Che,Dongsheng;Xu,Hao;Abbas,Abdulla;Wang,Bi-Cheng;Liu,Zhi-Jie
Many computational tools have been developed recently to accurately predict the structure of a protein from its amino acid sequence. 1–3 In general, if a query protein sequence shares at least 30% sequence identity with a protein sequence whose 3D structure has been determined, the structure of this query sequence can be modeled based on the template structure, using MODELLER software for example. 3, 4 Computational software, however, cannot guarantee accurate prediction for those new proteins that share low sequence similarity in PDB. Therefore, experimental methods such as X-ray crystallography and nuclear magnetic resonance (NMR) are still the main approaches for a protein structural study. 5, 6 The current target selection strategy of most structural genomics centers7–9 mainly focuses on the representatives of manually curated protein families (Pfam), 10–12 that is, the selected protein sequence shares at least one conserved domain with other members within a family. In this way, the solved representative structures can be used as structural templates to predict structures of the remaining protein sequences in the same family using computational tools. It has been shown that this ‘‘Pfam’’target selection strategy increases not only the number of novel structures, but also the number of new folds. 13 However, over-emphasis of Pfam and ignoring non-Pfam sequences (ie, not sharing any conserved domain in Pfam) in target selection might lead to biased distribution of Pfam and non-Pfam structures in PDB, and possibly slow the growth rate of new structures and folds. Our analysis on 150 microbial genomes showed that non-Pfam sequences account for 25–30% of all Open Reading Frames (ORFs) for most genomes, and some could reach to 60%(unpublished data). The high percentage of non-Pfams over all ORFs reminds us non-Pfams should not be neglected while devising a target selection strategy. On the other hand, these non-Pfam sequences for each genome are either paralogous non-Pfam (in which sequences have homologous partners within the same organism), or orthologous non-Pfam (in which sequences have orthologous partners in the closely