ECOD: Large scale classification of predicted and experimental protein structures
ECOD: Large scale classification of predicted and experimental protein structures
批准号:
10659763
负责人:
Richard Dustin Schaeffer
金额:
$34.44万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-01 至 2027-04-30
关键词:
3-DimensionalAccelerationAdoptedAmino Acid SequenceArtificial IntelligenceBase SequenceBenchmarkingBiochemicalBiologicalBiological AssayCatalogingClassificationCodeCollaborationsCommunitiesComputing MethodologiesDataData SetData SourcesDatabasesDedicationsDepositionDevelopmentDiseaseElectron MicroscopyEuropeanEvolutionFamilyFutureGeneticGrowthHomologous GeneHumanInfrastructureLaboratoriesLeadManualsMembraneMethodsModelingMolecular BiologyNMR SpectroscopyPathogenicityPathogenicity IslandPeptide Signal SequencesPeriodicalsPropertyProtein FamilyProteinsScienceSequence HomologySiteStructureSystemTechniquesTertiary Protein StructureTestingTimeUniversitiesUpdateVibrioVirulence FactorsWorkWorkloadX-Ray Crystallographycandidate identificationcaspase 14comparative genomicscomputational pipelinesdata qualitydatabase schemadeep learningexperimental studyhuman diseasehuman pathogenimprovedlearning networkmodel organismnovelpathogenpathogenic bacteriaprotein data bankprotein functionprotein structureprotein structure predictionstructural biologysuccessthree dimensional structuretoolweb interfaceweb pageweb site
中文摘要
项目摘要
蛋白质结构域的分类在历史上用于将由蛋白质结构域共同生成的3D结构数据上下文化。
实验结构测定方法,例如X射线晶体学,核磁共振光谱学,
和电子显微镜。我们的数据库,蛋白质结构域的进化分类(ECOD),已服务于生物学
七年来,社区从实验结构中编目域之间的进化关系。的
最近出现的高精度结构预测方法,如AlphaFold(AF)和RoseTTAFold(RF),以及
随后在AlphaFold数据库(AFDB)中发布了1000万个预测结构,预示着结构的范式转变。
生物学和领域分类。结构沉积的速度预计会在100到1000之间跳跃-
折我们建议利用这一革命,并将ECOD转化为一个全面的分类,
整个蛋白质大学使用序列,结构和功能的证据。通过同时分类实验
并预测了来自模式生物和人类病原体的蛋白质结构,我们的分类将有助于科学
社区批判性地评估结构模型,并利用进化信息来发现和实验
表征蛋白质功能。
对AF模型进行分类对ECOD管道提出了挑战,工作量增加50倍,
模型中的非球形和低质量区域。因此,我们的第一个目标是升级ECOD的基础设施,
从AF模型中识别单个结构域并整合序列、结构和功能位点相似性的方法
我们的自动分类。与目前依靠人类专家进行结构和
基于功能的分类,这些改进将大大减少手动管理的需要,并将允许我们
为了实现我们的第二个目标,即,通过以下组合将超过100万个已发布AF模型的域分类为ECOD
计算管道和最小的手动工作(0.25%-1%的情况)。利用洪水的自动对焦模型,新的
自动管道,以及人类策展人的专业知识,我们希望两者都能显着改善ECOD,并评估
通过(1)覆盖Pfam中所有已知的蛋白质家族,(2)通过进化确认远程同源性,
中间体,(3)比较进化相关的实验和预测的结构,以及(4)解决错误,
通过定期质量检查,最后,我们将率先进行功能性发现,
我们的第三个目标是研究细菌病原体中的毒力因子(VF),
或者由我们的实验合作者Orth实验室进行研究。快速发展的VF对于
通过序列进行结构预测或功能推断。我们将在24种细菌病原体中鉴定出候选的VF,
获得它们的结构模型,并利用与已知蛋白质在结构和功能位点上的相似性来推断它们的功能。
有希望的假设将在Orth实验室通过生化和遗传分析进行实验测试。
英文摘要
Project Summary
Classification of protein domains have historically served to contextualize the 3D structural data collectively generated by
experimental structure determination methods such as X-ray crystallography, nuclear magnetic resonance spectroscopy,
and electron microscopy. Our database, Evolutionary Classification of protein Domains (ECOD), has served the biological
community for seven years cataloguing evolutionary relationships between domains from experimental structures. The
recent advent of high-accuracy structure prediction methods, such as AlphaFold (AF) and RoseTTAFold (RF), and the
consequent release of 1 million predicted structures in AlphaFold Database (AFDB) heralds a paradigm shift in structural
biology and domain classification. The rate of structure deposition is expected to jump between a hundred to a thousand-
fold. We propose to take advantage of this revolution and transform ECOD into a comprehensive classification of the
entire protein university using sequence, structure, and functional evidence. By simultaneously classifying experimental
and predicted structures of proteins from model organisms and human pathogens, our classification will help the scientific
community to critically evaluate structure models and utilize the evolutionary information to discover and experimentally
characterize protein function.
Classifying AF models challenges the ECOD pipeline by a 50-fold increase in the workload and by the significant fraction of
non-globular and low-quality regions in the models. Thus, our first Aim is to upgrade ECOD’s infrastructure and develop
methods to identify single domains from AF models and to integrate sequence, structure, and functional site similarities
into our automatic classification. Compared to the current ECOD workflow that relies on human experts for structure-and-
function-based classification, these improvements will drastically decrease the need for manual curation and will allow us
to achieve our second Aim, i.e., classifying domains of over 1 million released AF models into ECOD via a combination of
computational pipelines and minimal manual efforts (0.25% 1% cases). Utilizing the deluge of AF models, the new
automatic pipeline, and expertise of human curators, we expect both to significantly improve ECOD and to evaluate the
quality of AF models by (1) covering all known protein families in Pfam, (2) confirming remote homology via evolutionary
intermediates, (3) comparing evolutionarily related experimental and predicted structures, and (4) resolving errors and
inconsistency through periodic quality checks. Finally, we will take the lead in making functional discoveries for
biomedically important proteins classified by ECOD in our third Aim, studying virulence factors (VFs) in bacterial pathogens
modelled by AFDB or studied by our experimental collaborators, the Orth lab. Fast evolving VFs were a challenge for
structure prediction or functional inference by sequence. We will identify candidate VFs in two dozen bacterial pathogens,
obtain their structure models, and infer their function using similarities to known proteins in structure and functional sites.
Promising hypotheses will be tested experimentally in the Orth lab through biochemical and genetic assays.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金