New Computational Methods for Data-driven Protein Structure Prediction
New Computational Methods for Data-driven Protein Structure Prediction
批准号:
8072026
负责人:
JINBO XU
金额:
$26.59万
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-05-14 至 2015-04-30
关键词:
AlgorithmsAmino Acid SequenceBiologicalBiological ProcessCategoriesCell CommunicationCellsComputer softwareComputing MethodologiesDNA SequenceDataDatabasesDecision TreesDevelopmentDiseaseDistantDrug IndustryGenerationsGenetic TranscriptionGoalsHealth Care CostsHomologous GeneHomology ModelingInternetKnowledgeLeadLearningLifeMachine LearningMaintenanceMetabolismMethodsModelingMolecular ConformationNon-linear ModelsPeptide Sequence DeterminationPharmaceutical PreparationsPharmacologic SubstancePlayPreventiveProcessProtein ConformationProteinsResearchRoleSamplingStagingStructureTechniquesTimeTranslationsTreesWorkbasedata modelingdatabase structuregenome sequencingimprovedknowledge baselecturesnovelnovel diagnosticsprotein foldingprotein structureprotein structure predictionpublic health relevancetherapeutic developmentthree dimensional structurewiki
中文摘要
描述(由申请人提供):蛋白质在所有生物过程中发挥核心作用。与基因组的完整测序类似,蛋白质结构的完整描述是了解生物生命的基本步骤,并且在治疗和药物开发方面也具有高度相关性。该项目的长期目标是通过两种独立但互补的策略,为数据驱动的蛋白质结构预测开发机器学习方法:1)对蛋白质数据库中具有远程同源物的蛋白质进行更准确的基于模板的建模;2)对没有可检测模板的蛋白质进行更好的无模板建模方法,并改进基于模板的模型。具体目标是:目标1)通过1a)使用基于回归树的非线性评分函数改善蛋白质序列-模板比对,特别是在没有良好序列谱的情况下,极大地改进基于模板的建模;1b)利用结合残差级和原子级特征的机器学习方法改进折叠识别;目的2)通过三种独立但互补的方法改进连续空间中的蛋白质构象采样,从而实现无模板建模:2a)使用条件(马尔可夫)随机场(CRF)模型建模非线性序列-结构关系;2b)同时采样二级和三级结构;2c)从模板中学习结构信息。该项目的核心是通过从现有的序列/结构数据库中学习蛋白质序列-结构关系,开发各种CRF模型,用于数据驱动的蛋白质结构预测。这项研究的成果包括一个基于回归树的CRF模型,用于精确的蛋白质比对,特别是对于PDB中没有密切同源物或没有非常好的序列谱的蛋白质;蛋白质折叠识别的支持向量机模型;连续空间中高效蛋白质构象采样的CRF模型;并有完整的蛋白质结构预测软件包。此外,它将为各种学术和生物医学用户提供一个公开的web服务器。蛋白质结构预测将导致广泛的生物医学应用,例如开发新的诊断方法,更好地了解疾病过程和改进预防疗法,从而降低医疗保健费用。蛋白质建模也广泛应用于制药行业,并融入到制药研究的大多数阶段。
英文摘要
DESCRIPTION (provided by applicant): Proteins play a central role in all biological processes. Akin to the complete sequencing of genomes, complete description of protein structures is a fundamental step towards understanding biological life, and is also highly relevant medically in the development of therapeutics and drugs. The broad, long-term goal of the project is to develop machine learning methods for data-driven protein structure prediction through two independent but complementary strategies: 1) much more accurate template-based modeling for proteins with remote homologs in the Protein Data Bank and 2) better template-free modeling method for proteins without detectable templates and for improving template-based models. The specific aims are: Aim 1) to greatly improve template-based modeling by 1a) improving protein sequence-template alignment using a regression-tree-based nonlinear scoring function, especially when good sequence profiles are unavailable; and 1b) improving fold recognition using a machine learning method to combine both residue-level and atom-level features; Aim 2) to improve protein conformation sampling in a continuous space and thus template-free modeling by three independent but complementary approaches: 2a) modeling nonlinear sequence- structure relationship using Conditional (Markov) Random Fields (CRF) models; 2b) simultaneously sampling secondary and tertiary structure; and 2c) learning structure information from template. The core of the project is to develop various CRF models for data-driven protein structure prediction, by learning protein sequence-structure relationship from existing sequence/structure databases. The product of this research includes a regression-tree-based CRF model for accurate protein alignment, especially for proteins without close homologs in the PDB or without very good sequence profiles; a SVM model for protein fold recognition; a few CRF models for efficient protein conformation sampling in a continuous space; and a complete protein structure prediction software package. Also, it will produce a web server publicly available for various academic and biomedical users. Protein structure prediction will lead to a broad range of biomedical applications, such as the development of novel diagnostics, better understanding of disease processes and improved preventive therapies leading to reduced health care costs. Protein modeling is also widely applied in the pharmaceutical industry and integrated into most stages of pharmaceutical research.
PUBLIC HEALTH RELEVANCE:
Novel protein structure prediction will lead to a broad range of biomedical applications, such as the development of novel diagnostics, better understanding of disease processes and improved preventive therapies leading to reduced health care costs. Protein modeling is also widely applied in the pharmaceutical industry and integrated into most stages of pharmaceutical research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10693195
-
项目类别:
-
资助金额:$32.66万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8657055
-
项目类别:
-
资助金额:$26.59万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8269822
-
项目类别:
-
资助金额:$26.59万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:7764110
-
项目类别:
-
资助金额:$26.86万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10477350
-
项目类别:
-
资助金额:$32.4万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:8463561
-
项目类别:
-
资助金额:$25.66万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
New Computational Methods for Data-driven Protein Structure Prediction
-
批准号:10246779
-
项目类别:
-
资助金额:$32.16万
-
财政年份:2010
-
负责人:JINBO XU
-
依托单位:
A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE
-
批准号:7723215
-
项目类别:
-
资助金额:$0.05万
-
财政年份:2008
-
负责人:JINBO XU
-
依托单位:
A STUDY OF THE INTEROPERABILITY BETWEEN TERAGRID AND CNGRID BY EXPERIMENTING RE
-
批准号:7601478
-
项目类别:
-
资助金额:$0.03万
-
财政年份:2007
-
负责人:JINBO XU
-
依托单位:
海外基金