课题基金 / 基金详情

UKRI/BBSRC-NSF/BIO: Unifying Pfam protein sequence and ECOD structural classifications with structure models

UKRI/BBSRC-NSF/BIO: Unifying Pfam protein sequence and ECOD structural classifications with structure models
UKRI/BBSRC-NSF/BIO:通过结构模型统一 Pfam 蛋白质序列和 ECOD 结构分类
批准号:
2224128
负责人:
Nick Grishin
金额:
$114.63万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-11-01 至 2025-10-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
蛋白质是复杂的有机分子,对生命的所有功能都至关重要。对蛋白质的研究始于对其进化亲属的了解。蛋白质在进化过程中倾向于保持其功能。因此,一个更深入研究的亲属可以让研究人员了解一个不太了解的亲属的特性,从而提出有待检验的假设。一种蛋白质的近亲被编目在公共的、免费的分类数据库中。虽然大多数蛋白质的氨基酸序列是已知的,但只有一小部分蛋白质的空间结构是通过实验确定的。蛋白质分类既可以基于蛋白质序列(包括那些未知的3D结构),也可以基于蛋白质空间结构(增加序列)。基于结构的分类更准确,但包含的蛋白质较少,必然会遗漏那些结构未知的蛋白质。最近开发了革命性的结构预测方法,可以为任何蛋白质序列产生准确的3D模型,桥接这些序列和结构/序列分类。现在是通过协同合作将两种分类类型结合在一起的时候了。美国(ECOD:蛋白质结构域进化分类数据库,主要基于结构)和英国(Pfam:蛋白质家族数据库,主要基于序列)的研究小组将共同努力,使他们的两个数据库相互一致,更加准确,以造福科学家和更广泛的社区。这些蛋白质分类的结果很容易被纳入许多其他资源,如维基百科页面,因此被科学和教育领域的受众广泛使用。ECOD和Pfam团队将合作对AlphaFold和RoseTTAfold生成的100多万个蛋白质结构模型进行分类,将其划分为进化近亲的蛋白质家族。现有的家族将增加额外的蛋白质。新的家族将通过序列剖面相似性来确定,并辅以结构分析,以符合Pfam分类标准的方式确定。该项目需要升级软件基础设施,以处理数百万个模型,并根据域标识符和家族名称同步两种分类。ECOD和Pfam分类的同步将通过四种方式实现:1)在Pfam中为目前仅存在于ECOD的领域定义新的科;2)通过修正ECOD分类,使所有域都归类为Pfam族(必要时定义新的域);3)将包含多个ECOD域的Pfam域家族拆分为包含单个域的多个Pfam域家族。4)通过在两个数据库中使用具有3D模型的蛋白质,对结构域边界和家族分类做出一致的协作决策。这些一致的领域定义和分类将有助于在科学界和广大公众中广泛产生功能推断和检测进化见解。最后,互联网架构将升级,通过门户网站向更广泛的科学界提供这些领域的数据。该项目的成果可以在ECOD http://prodata.swmed.edu/ecod和Pfam http://pfam.xfam.org.This奖项中找到,这反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Proteins are complex organic molecules essential for all functions of life. A study of a protein begins from learning about its evolutionary relatives. Proteins tend to retain their function as they evolve. Therefore, a more closely studied relative can inform researchers about the properties of a less understood relative, suggesting hypotheses to be tested. A protein’s relatives are cataloged in public, freely available, classification databases. Although amino-acid sequences are known for most proteins, spatial structures have only been experimentally determined for a small fraction. Protein classification can be based on either protein sequences (including those with yet unknown 3D structure), or protein spatial structures (augmented with sequences). Structure-based classifications are more accurate but include fewer proteins, necessarily missing those whose structures are unknown. Recently developed revolutionary structure prediction methods that can produce accurate 3D models for any protein sequence bridge these sequence-only and structure/sequence classifications. Now is the time to bring the two classification types together through a synergistic collaboration. The teams in the United States (ECOD: Evolutionary Classification of protein Domains database, mostly structure-based) and the United Kingdom (Pfam: Protein families database, mostly sequence-based) will work together to make their two databases consistent with each other and more accurate for the benefit of scientists and the broader community. The results of these protein classifications are readily incorporated into many other resources, such as Wikipedia pages, and thus are widely used by an audience in science and education.The ECOD and Pfam teams will collaboratively classify more than 1 million protein structure models generated by AlphaFold and RoseTTAfold into protein families of close evolutionary relatives. Existing families will be expanded with additional proteins. New families will be defined by sequence profile similarity aided by structure analysis in a manner consistent with the Pfam classification standards. The project requires upgrading the software infrastructure to process millions of models and to synchronize the two classifications in terms of domain identifiers and family names. Synchronization of the ECOD and Pfam classifications will be achieved in four ways: 1) By defining new families in Pfam for domains currently present only in ECOD; 2) By rectifying the ECOD classification such that all domains are classified into a Pfam family (defining new ones where necessary); 3) By splitting Pfam domain families containing multiple ECOD domains into multiple families containing single domains. 4) By making consistent collaborative decisions about domain boundaries and family classification using proteins with 3D models in both databases. These consistent domain definitions and classifications will facilitate broad generation of functional inference and detection of evolutionary insights in the scientific community and the public at large. Lastly, the internet architecture will be upgraded to serve these domain data to the broader scientific community through web portals. The results of this project can be found incorporated into both ECOD http://prodata.swmed.edu/ecod and Pfam http://pfam.xfam.org.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DPAM : A domain parser for AlphaFold models
DPAM:AlphaFold 模型的域解析器
DOI: 10.1002/pro.4548
发表时间: 2023
期刊: Protein Science
影响因子: 8
作者: [Zhang, Jing, Schaeffer, R. Dustin, Durham, Jesse, Cong, Qian, Grishin, Nick V.]
通讯作者: Grishin, Nick V.
海外基金