Evolutionary and thermodynamical features of musculoskeletal disease mutations is human intrinsically disordered protein regions
Evolutionary and thermodynamical features of musculoskeletal disease mutations is human intrinsically disordered protein regions
批准号:
2114913
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
有些蛋白质的三维结构特征很好。特别是,已知它们的一部分充当“功能”单元,例如活性和结合位点或功能域。然而,在蛋白质内部,存在内在无序的区域,这些区域的特征是缺乏明确的三维结构。尽管这些无序区域没有显示出特定的高构象状态,但已知它们在功能上很重要,例如通过它们参与蛋白质-蛋白质相互作用或DNA/RNA结合。在最近的一项工作中,我们已经表明,存在持续的积极选择,这对人类长期内在无序的蛋白质区域的进化做出了重大贡献。此外,这些蛋白质区域富含翻译后修饰位点以及区域和基序(具有生物学重要性的注释序列延伸),但令人惊讶的是,除了与肌肉骨骼疾病相关的突变外,疾病突变往往在紊乱区域发生的频率要低得多(Uversky等人,2014)。这个及时的项目旨在了解为什么疾病突变在无序的蛋白质区域往往不太频繁。这个项目的重点将放在肌肉骨骼疾病中异常的突变组上,这些突变组富集于无序的蛋白质区域,通过将它们与其他疾病组中的突变组进行比较。基础将是一种新的基因组比较方法,目前在Gossmann实验室开发的目标是识别疾病相关位点的进化特性。此外,我们将与利物浦大学的Daniel Rigden合作,进行分子动力学模拟,以研究蛋白质柔韧性和蛋白质-配体相互作用的疾病相关突变的三维特征。此外,我们将利用机器学习方法来预测单位点残基效应上的蛋白质紊乱。在与谢菲尔德大学Mirre Simons实验室的合作下,实验证据和疾病候选位点的潜在机制可以在果蝇模型中进行功能测试。这将最终深入了解疾病突变是否真的不太可能发生在无序蛋白质区域,或者我们对与无序蛋白质区域相关的疾病特性的缺乏了解是否导致了相应数据库中的注释不足。对于这个高度创新,协作和跨学科的博士学位,相应的候选人应该在生物学,分子生物学和遗传学以及基本的编程知识方面有很强的背景,或者至少对研究基本生物学问题的计算方法有浓厚的兴趣。在生物信息学背景是有利的,但是所有必要的方法将在博士期间教授。该项目将利用多个生物学“大”数据集,如1000基因组计划、Uniprot、大型哺乳动物系统发育及其全基因组信息,以及PDB和多个二级数据库。
英文摘要
There are proteins that are well characterized with regard to their three dimensional structure. In particular it is known that parts of them act as "functional" units, e.g. active and binding sites or functional domains. However, within proteins, regions of intrinsically disorder occur and these are characterized by a lack of a well-defined three-dimensional structure. Although these disordered regions do not show a particular higher conformational state, they are known to be functionally important, such as through their involvement in protein-protein interactions or DNA/RNA binding. In a recent work we have shown that there is ongoing positive selection that contributes substantially to the evolution of human long intrinsically disordered protein regions. Furthermore, these protein regions are enriched in posttranslational modification sites as well as regions and motifs (annotated sequence stretches of biological importance), but surprisingly disease mutations tend to occur much less frequently in disordered regions (Uversky et al., 2014), with the exception of mutations associated with musculoskeletal diseases. This timely project aims to understand why disease mutations tend to be less frequent in disordered protein regions. The focus of this project will lie on the exceptional group of mutations involved in musculoskeletal diseases that are enriched in disordered protein regions by comparing them to those involved in other disease groups. Fundamental will be a novel genomic comparative approach currently developed in the Gossmann lab targeted at the identification of the evolutionary properties of disease-associated sites. Furthermore, in collaboration with Daniel Rigden from the University of Liverpool, we will conduct molecular dynamic simulations to investigate three dimensional features of disease associated mutations on protein flexibility and protein-ligand interactions. Furthermore we will exploit machine learning approaches to predict protein disorder on the single site residue effects. Experimental evidence and the underlying mechanistics of disease candidate sites can then functionally be tested in a fly model in collaboration with the Mirre Simons lab at the University of Sheffield. This will ultimately gain insights into whether disease mutations are genuinely less likely to occur in disordered protein regions or whether our lacking understanding of disease properties associated with disordered protein regions has led to an under-annotation in the respective databases. For this highly innovative, collaborative and interdisciplinary PhD the respective candidate should have a strong background in biology, molecular biology and genetics as well basic programming knowledge or at least a strong interest in computational approaches to investigate fundamental biological problems. A background in bioinformatics is of advantage, however all necessary approaches will be taught during the duration of the PhD. This project will take advantage of multiple biological "big" data sets, such as the 1000-Genome project, Uniprot, large-scale mammalian phylogenies and their respective whole genome information, as well as PDB and several secondary databases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金