Development of an accurate classification system of proteins into structured and unstructured regions that uncovers novel structural domains: its application to human transcription factors

Development of an accurate classification system of proteins into structured and unstructured regions that uncovers novel structural domains: its application to human transcription factors
复制标题

DOI:
10.1186/1472-6807-9-26
复制
发表时间:
2009-04-30
影响因子:
--
通讯作者:
Nishikawa, Ken
Nishikawa, Ken
中科院分区:
生物4区
文献类型:
--
作者:
Fukuchi, Satoshi;Homma, Keiichi;Nishikawa, Ken

文献摘要

被引文献

相似文献

背景:除了结构域,大多数真核生物蛋白质具有内在无序(ID)区域。虽然ID区域通常发挥重要的功能作用,但它们的准确识别是困难的。由于人类转录因子是一类典型的具有长ID区的蛋白质,我们将其作为所有蛋白质的模型,并试图将转录因子准确地分为结构域和ID区。虽然在我们以前的研究中在人TF中检测到除了DNA结合和/或其他结构域之外的极高比例的ID区域,但仍有20%的残基未分配。在这份报告中,我们利用ID区域中的序列差异通常比结构区域中的更高,将蛋白质完全分为结构域和ID区域。新的二分系统首先识别已知结构的结构域,然后结合现有的工具和新开发的基于序列差异的程序来分配结构域和ID区域,考虑到不结盟地区。该系统被发现是高度准确的:其应用于一组蛋白质与实验验证的ID区域的错误率低至2%。将该系统应用于人TF(401个蛋白质)表明,38%的残基在结构域中,而62%在ID区域中。ID区域的优势与大肠杆菌的TF(229个蛋白质)形成鲜明对比,其中只有5%落在ID区域。该方法还显示,人和E.结论:本系统验证了包含未比对区域信息的序列差异是识别ID区域的良好指标。该系统首次估计了人类TF中结构化/非结构化区域的完整分数,还揭示了与已知结构没有同源性的结构域。这些预测的新结构域是结构基因组学的良好靶点。当应用于其他蛋白质时,该系统有望发现更多新的结构域。
Background: In addition to structural domains, most eukaryotic proteins possess intrinsically disordered (ID) regions. Although ID regions often play important functional roles, their accurate identification is difficult. As human transcription factors (TFs) constitute a typical group of proteins with long ID regions, we regarded them as a model of all proteins and attempted to accurately classify TFs into structural domains and ID regions. Although an extremely high fraction of ID regions besides DNA binding and/or other domains was detected in human TFs in our previous investigation, 20% of the residues were left unassigned. In this report, we exploit the generally higher sequence divergence in ID regions than in structural regions to completely divide proteins into structural domains and ID regions.Results: The new dichotomic system first identifies domains of known structures, followed by assignment of structural domains and ID regions with a combination of pre-existing tools and a newly developed program based on sequence divergence, taking un-aligned regions into consideration. The system was found to be highly accurate: its application to a set of proteins with experimentally verified ID regions had an error rate as low as 2%. Application of this system to human TFs (401 proteins) showed that 38% of the residues were in structural domains, while 62% were in ID regions. The preponderance of ID regions makes a sharp contrast to TFs of Escherichia coli (229 proteins), in which only 5% fell in ID regions. The method also revealed that 4.0% and 11.8% of the total length in human and E. coli TFs, respectively, are comprised of structural domains whose structures have not been determined.Conclusion: The present system verifies that sequence divergence including information of unaligned regions is a good indicator of ID regions. The system for the first time estimates the complete fractioning of structured/un-structured regions in human TFs, also revealing structural domains without homology to known structures. These predicted novel structural domains are good targets of structural genomics. When applied to other proteins, the system is expected to uncover more novel structural domains.