Identification and mapping of self-assembling protein domains encoded by the Escherichia coli K-12 genome by use of λ repressor fusions

Identification and mapping of self-assembling protein domains encoded by the Escherichia coli K-12 genome by use of λ repressor fusions
复制标题

DOI:
10.1128/jb.186.5.1311-1319.2004
复制
发表时间:
2004-03-01
影响因子:
3.2
通讯作者:
Hu, JC
Hu, JC
中科院分区:
生物学3区
文献类型:
--
作者:
Mariño-Ramírez, L;Minor, JL;Hu, JC

文献摘要

被引文献

相似文献

从大肠杆菌K-12菌株MG1655中鉴定出大肠杆菌基因组编码的自组装蛋白和蛋白片段。将随机DNA片段克隆到一系列lambda抑制子融合载体中,进行抗Lambda噬菌体感染的筛选。通过对插入片段末端进行测序来鉴定幸存者,并从已知的基因组序列推断融合蛋白序列。从2,089个候选序列中回收了463个非冗余开放阅读框架编码的相互作用序列标签(IST)。这些IST的长度从16到794个氨基酸不等,被聚集成重叠片段家族,识别由232个大肠杆菌基因编码的潜在同型相互作用。阻遏物融合发现了来自每个基于蛋白质的功能类别的基因的IST,但膜蛋白的表达不足。含有IST的基因富含调节蛋白和形成高阶低聚物的蛋白质。ISTS鉴定的48个同型蛋白(20.7%)预测含有卷曲的卷曲。虽然大多数含有IST的基因与其他细菌基因组中的蛋白质有明显的相关性,但超过一半的IST基因在蛋白质数据库中没有可识别的同源基因,这表明它们可能包括许多新的结构。
Self-assembling proteins and protein fragments encoded by the Escherichia coli genome were identified from E. coli K-12 strain MG1655. Libraries of random DNA fragments cloned into a series of lambda repressor fusion vectors were subjected to selection for immunity to infection by phage lambda. Survivors were identified by sequencing the ends of the inserts, and the fused protein sequence was inferred from the known genomic sequence. Four hundred sixty-three nonredundant open reading frame-encoded interacting sequence tags (ISTs) were recovered from sequencing 2,089 candidates. These ISTs, which range from 16 to 794 amino acids in length, were clustered into families of overlapping fragments, identifying potential homotypic interactions encoded by 232 E. coli genes. Repressor fusions identified ISTs from genes in every protein-based functional category, but membrane proteins were underrepresented. The IST-containing genes were enriched for regulatory proteins and for proteins that form higher-order oligomers. Forty-eight (20.7%) homotypic proteins identified by ISTs are predicted to contain coiled coils. Although most of the IST-containing genes are identifiably related to proteins in other bacterial genomes, more than half of the ISTs do not have identifiable homologs in the Protein Data Bank, suggesting that they may include many novel structures.