Modeling evolution of functional context in proteins
Modeling evolution of functional context in proteins
批准号:
8055568
负责人:
DAVID D POLLOCK
金额:
$34.55万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-03-01 至 2013-02-28
关键词:
AffectAgingAlzheimer&aposs DiseaseAmino Acid SequenceAmino AcidsApoptosisApoptoticAreaAttentionBiodiversityBiologicalCell DeathCell physiologyComplexComputing MethodologiesDNADataData SetDependenceDevelopmentDiabetes MellitusDiseaseDrug DesignEnvironmentEtiologyEvaluationEvolutionFunctional disorderGenesGenetic VariationGenomeGenomicsHealthHumanHuman BiologyIndividualLeadLifeMalignant NeoplasmsMethodsMitochondriaMitochondrial ProteinsModelingMutationNuclearOrganismOxidative PhosphorylationParkinson DiseasePathologyPathway interactionsPatternPeptide Sequence DeterminationPhylogenyPhysiologicalPlayPositioning AttributePrimatesProbabilityProcessProteinsResearchRoleSamplingSiteSpecific qualifier valueStructureTestingThermodynamicsTimeTranslational ResearchUrsidae FamilyVariantWorkage relatedbasecomparativecomparative genomicsdesignexperiencefitnessflexibilityhuman diseaseimprovedinsightmitochondrial genomenervous system disordernovelpreventprotein functionprotein structure functionprotein structure predictionpublic health relevancereconstructionresponse
中文摘要
描述(由申请人提供):随着蛋白质随着时间的推移积累变化和分化,它们必须继续满足使它们能够正常发挥功能的结构和能量限制。正因为如此,生物体中蛋白质序列的多样性代表了蛋白质序列、结构和功能之间关系的丰富数据。然而,从这些数据中提取洞察力仍然是一个挑战。原则上,解码蛋白质序列生物多样性中的功能信息的最佳方法是使用基于系统发育的现实模型的参数统计推断。从历史上看,这种方法一直受到可用序列数据量的限制,以及所需概率计算的困难和令人望而却步的计算复杂性。在后基因组时代,序列生物多样性已经变得更加容易获得,然而,尽管有几个小组取得了重要的进展,但绝大多数利用比较序列信息的研究依赖于带有许多出于便利的假设的模型,这些模型忽略了重要的生物学效应,例如蛋白质中残基的结构和功能上下文的变化,以及随着时间的推移。这些简化阻碍了比较基因组数据在健康和疾病中发挥基因变异作用的全部潜力,因此可能对生物和生物医学的发现构成重大障碍。在这里,建议利用最近开发的方法,我们已经设计了这些方法来通过消除潜在的误导性模型过度简化来纠正这种情况。我们将使用这些方法来模拟蛋白质中不同位置的进化模式中复杂的上下文相关变化,并模拟这些模式如何随着时间的推移和对外部影响的反应而变化。不同位置的进化模式将与已知的结构和能量特征有关,在这些分析的基础上,将开发直接结合真实生物效应的新模型。由于我们的算法进步,这些模型可以更准确地反映进化过程中相互依赖的序列、结构和功能影响的真实复杂性,产生无与伦比的能力来检测微妙但有意义的影响,提供更好的预测能力,并更准确地表征蛋白质进化的原因和结果。这项拟议的研究对人类健康具有广泛的重要意义,因为蛋白质功能和功能障碍在绝大多数人类疾病和健康障碍的机制和病因中发挥着核心作用。因此,更好地了解序列变异、结构和功能之间的关系将更好地预测人类突变的影响,更好地了解蛋白质功能及其在人类生物学中的作用,改进可能用于改善人类健康的新蛋白质的合理设计,并有可能更准确地进行基于结构的药物设计。拟议的研究将专注于对完整脊椎动物线粒体基因组中编码的蛋白质进行建模,利用不同脊椎动物物种中独特的密集序列采样。我们将详细关注灵长类动物,包括人类和线粒体基因组,以便我们的研究将提供具体的洞察力,了解关键的氧化磷酸化蛋白是如何发挥作用的,以及这些基因的突变如何通过破坏结构和功能而导致人类疾病。鉴于线粒体对衰老、疾病(如糖尿病、帕金森氏症、阿尔茨海默氏症和其他神经疾病)以及包括细胞凋亡在内的细胞过程的核心重要性,这些见解可能被证明是直接有益的,作为几个领域的实验、药理学和翻译研究的假设生成和测试平台。通过设计,该项目将为未来对核基因组数据集的研究铺平道路,因为它们的数量和多样性都在增加。
英文摘要
DESCRIPTION (provided by applicant): As proteins accumulate change and diverge over time, they must continue to satisfy the structural and energetic constraints that enable them to function properly. Because of this, the diversity of protein sequences available from living organisms represents a wealth of data on the relationships between protein sequence, structure, and function. Extracting insights from these data, however, remains a challenge. In principle, the optimal approach to decoding functional information in protein sequence biodiversity is to use parametric statistical inference with realistic phylogeny-based models. Historically, this approach has been limited by the amount of sequence data available, and by the difficulty and prohibitive computational complexity of the probability calculations needed. Sequence biodiversity has become much more readily available in the post-genomic era, and yet despite important advances by a few groups, the vast majority of research making use of comparative sequence information has depended on models with numerous convenience-motivated assumptions that ignore important biological effects such as the variation of structural and functional contexts across residues in a protein, and over time. These simplifications prevent the full potential of comparative genomic data from being brought to bear on the role of genetic variation in health and disease, and thus may pose significant roadblocks to biological and biomedical discovery. Here, it is proposed to take advantage of recently developed methods, which we have designed to remedy this situation by eliminating the need for potentially misleading model over-simplifications. We will use these methods to model complex context-dependent variation in evolutionary patterns across positions in proteins, and to model how these patterns change over time and in response to external influences. The evolutionary patterns at different positions will be related to known structural and energetic features and, based on these analyses, new models will be developed that directly incorporate bona fide biological effects. Because of our algorithmic advances, these models can be built to more accurately reflect the true complexity of interdependent sequence, structural and functional effects on evolutionary processes, yielding unparalleled power to detect subtle but meaningful effects, to provide better predictive capabilities, and to more precisely characterize the causes and consequences of protein evolution. The proposed research is broadly important for human health because of the central role that protein function and dysfunction play in the mechanisms and etiology of a vast majority of human diseases and health disorders. Thus, a better understanding of the relationship between sequence variation, structure, and function will yield better prediction of the effects of human mutations, greater understanding of protein function and its role in human biology, improved rational design of novel proteins that might be used to improve human health, and potentially more accurate structure-based drug design. The proposed research will focus on modeling proteins encoded in complete vertebrate mitochondrial genomes, taking advantage of the uniquely dense sequence sampling available across diverse vertebrate species. We will pay detailed attention to primate, including human, mitochondrial genomes, so that our research will provide specific insight into how key oxidative phosphorylation proteins function, and how mutations in these genes lead to human diseases by disrupting structure and function. Given the central importance of the mitochondrion to aging, disease (e.g., diabetes, Parkinson's, Alzheimer's, and other neurological diseases) and to cellular processes including apoptotic cell death, such insights may prove directly beneficial as a hypothesis-generating and testing platform for experimental, pharmacological, and translational research in several areas. By design, this project will pave the way for future research on nuclear genome datasets as their number and diversity increases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Genome-wide mutation models to decipher function
-
批准号:8776584
-
项目类别:
-
资助金额:$5.83万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Genome-wide mutation models to decipher function
-
批准号:9005906
-
项目类别:
-
资助金额:$4.0万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Genome-wide mutation models to decipher function
-
批准号:8606470
-
项目类别:
-
资助金额:$27.39万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Genome-wide mutation models to decipher function
-
批准号:8454425
-
项目类别:
-
资助金额:$26.13万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Genome-wide mutation models to decipher function
-
批准号:8238861
-
项目类别:
-
资助金额:$26.75万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Genome-wide mutation models to decipher function
-
批准号:8803386
-
项目类别:
-
资助金额:$27.81万
-
财政年份:2012
-
负责人:DAVID D POLLOCK
-
依托单位:
Modeling evolution of functional context in proteins
-
批准号:8143199
-
项目类别:
-
资助金额:$2.32万
-
财政年份:2009
-
负责人:DAVID D POLLOCK
-
依托单位:
Modeling evolution of functional context in proteins
-
批准号:9262235
-
项目类别:
-
资助金额:$35.77万
-
财政年份:2009
-
负责人:DAVID D POLLOCK
-
依托单位:
Modeling evolution of functional context in proteins
-
批准号:7767710
-
项目类别:
-
资助金额:$30.17万
-
财政年份:2009
-
负责人:DAVID D POLLOCK
-
依托单位:
Modeling evolution of functional context in proteins
-
批准号:8245205
-
项目类别:
-
资助金额:$32.12万
-
财政年份:2009
-
负责人:DAVID D POLLOCK
-
依托单位:
Protein sequence, structure, and computational analysis
-
批准号:7269040
-
项目类别:
-
资助金额:$23.11万
-
财政年份:2002
-
负责人:DAVID D POLLOCK
-
依托单位:
Protein sequence, structure, and computational analysis
-
批准号:6834875
-
项目类别:
-
资助金额:$22.22万
-
财政年份:2002
-
负责人:DAVID D POLLOCK
-
依托单位:
Protein sequence, structure, and computational analysis
-
批准号:6480118
-
项目类别:
-
资助金额:$13.83万
-
财政年份:2002
-
负责人:DAVID D POLLOCK
-
依托单位:
Protein sequence, structure, and computational analysis
-
批准号:6924665
-
项目类别:
-
资助金额:$2.11万
-
财政年份:2002
-
负责人:DAVID D POLLOCK
-
依托单位:
Protein sequence, structure, and computational analysis
-
批准号:6630490
-
项目类别:
-
资助金额:$13.82万
-
财政年份:2002
-
负责人:DAVID D POLLOCK
-
依托单位:
海外基金