New Computational Methods for Data-driven Protein Structure Prediction
New Computational Methods for Data-driven Protein Structure Prediction
批准号:
8657055
负责人:
JINBO XU
金额:
$26.59万
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-05-14 至 2015-09-24
关键词:
AlgorithmsAmino Acid SequenceBiologicalBiological ProcessCategoriesCell CommunicationCellsComputer softwareComputing MethodologiesDNA SequenceDataDatabasesDecision TreesDevelopmentDiseaseDistantDrug IndustryGenerationsGenetic TranscriptionGoalsHealth Care CostsHomologous GeneHomology ModelingInternetKnowledgeLeadLearningLifeMachine LearningMaintenanceMetabolismMethodsModelingMolecular ConformationNon-linear ModelsPeptide Sequence DeterminationPharmaceutical PreparationsPharmacologic SubstancePlayPreventiveProcessProtein ConformationProteinsResearchRoleSamplingStagingStructureTechniquesTimeTranslationsTreesWorkbasedata modelingdatabase structuregenome sequencingimprovedknowledge baselecturesnovelnovel diagnosticsprotein foldingprotein structureprotein structure predictionpublic health relevancetherapeutic developmentthree dimensional structurewiki
中文摘要
描述(由申请人提供):蛋白质在所有生物过程中起着核心作用。与基因组的完整测序类似,完整的蛋白质结构描述是理解生物生命的基本步骤,在治疗和药物的开发中也具有高度的医学相关性。该项目的广泛和长期目标是通过两种独立但互补的策略开发用于数据驱动的蛋白质结构预测的机器学习方法:1)对蛋白质数据库中具有远程同源的蛋白质进行更准确的基于模板的建模;2)对没有可检测到模板的蛋白质进行更好的无模板建模方法和改进基于模板的模型。其具体目标是:目标1)通过以下方式极大地改进基于模板的建模:1a)使用基于回归树的非线性评分函数改进蛋白质序列模板比对,特别是在无法获得好的序列图谱的情况下;以及1b)使用结合残基水平和原子水平特征的机器学习方法来改进折叠识别;目标2)改进连续空间中的蛋白质构象采样,从而通过三种独立但互补的方法来进行无模板建模:2a)使用条件(马尔可夫)随机场(CRF)模型来建模非线性序列-结构关系;2b)同时采样二级和三级结构;以及2c)从模板学习结构信息。该项目的核心是通过从现有的序列/结构数据库中学习蛋白质序列-结构关系,开发用于数据驱动的蛋白质结构预测的各种CRF模型。这项研究的产品包括:基于回归树的CRF模型,用于精确的蛋白质比对,特别是对于在PDB中没有紧密同源或没有非常好的序列图谱的蛋白质;用于蛋白质折叠识别的支持向量机模型;几个用于在连续空间中进行高效蛋白质构象采样的CRF模型;以及完整的蛋白质结构预测软件包。此外,它还将生产一个网络服务器,供各种学术和生物医学用户公开使用。蛋白质结构预测将导致广泛的生物医学应用,例如开发新的诊断方法,更好地了解疾病过程,改进预防疗法,从而降低医疗保健成本。蛋白质建模也被广泛应用于制药行业,并被整合到制药研究的大多数阶段。
英文摘要
DESCRIPTION (provided by applicant): Proteins play a central role in all biological processes. Akin to the complete sequencing of genomes, complete description of protein structures is a fundamental step towards understanding biological life, and is also highly relevant medically in the development of therapeutics and drugs. The broad, long-term goal of the project is to develop machine learning methods for data-driven protein structure prediction through two independent but complementary strategies: 1) much more accurate template-based modeling for proteins with remote homologs in the Protein Data Bank and 2) better template-free modeling method for proteins without detectable templates and for improving template-based models. The specific aims are: Aim 1) to greatly improve template-based modeling by 1a) improving protein sequence-template alignment using a regression-tree-based nonlinear scoring function, especially when good sequence profiles are unavailable; and 1b) improving fold recognition using a machine learning method to combine both residue-level and atom-level features; Aim 2) to improve protein conformation sampling in a continuous space and thus template-free modeling by three independent but complementary approaches: 2a) modeling nonlinear sequence- structure relationship using Conditional (Markov) Random Fields (CRF) models; 2b) simultaneously sampling secondary and tertiary structure; and 2c) learning structure information from template. The core of the project is to develop various CRF models for data-driven protein structure prediction, by learning protein sequence-structure relationship from existing sequence/structure databases. The product of this research includes a regression-tree-based CRF model for accurate protein alignment, especially for proteins without close homologs in the PDB or without very good sequence profiles; a SVM model for protein fold recognition; a few CRF models for efficient protein conformation sampling in a continuous space; and a complete protein structure prediction software package. Also, it will produce a web server publicly available for various academic and biomedical users. Protein structure prediction will lead to a broad range of biomedical applications, such as the development of novel diagnostics, better understanding of disease processes and improved preventive therapies leading to reduced health care costs. Protein modeling is also widely applied in the pharmaceutical industry and integrated into most stages of pharmaceutical research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10693195
-
项目类别:
-
资助金额:$32.66万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8269822
-
项目类别:
-
资助金额:$26.59万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:7764110
-
项目类别:
-
资助金额:$26.86万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10477350
-
项目类别:
-
资助金额:$32.4万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8463561
-
项目类别:
-
资助金额:$25.66万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8072026
-
项目类别:
-
资助金额:$26.59万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10246779
-
项目类别:
-
资助金额:$32.16万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE
-
批准号:7723215
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:JINBO XU
-
依托单位:
A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE
-
批准号:7601478
-
项目类别:
-
资助金额:$0.03万
-
财政年份:2007
-
负责人:JINBO XU
-
依托单位:
海外基金