Prediction of specificity determining sites in proteins from human population genetic variation by deep learning
Prediction of specificity determining sites in proteins from human population genetic variation by deep learning
批准号:
2468759
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
20多年来,我们的研究一直致力于开发有效的计算方法,从氨基酸序列预测蛋白质的功能、结构和特异性。这一经验被封装在BBSRC资助的广泛使用的软件工具中,其中包括Jalview(www.jalview.org)序列分析工作台和JPred(www.CompBio.dundee.ac.uk/jpred),Jalview(www.jalview.org)在全球拥有70,000多名固定用户,JPred(www.CompBio.dundee.ac.uk/jpred)每月为英国和国际实验室的科学家执行多达250,000个二级结构和来自氨基酸序列的其他特征的预测。Jalview和JPred总共被描述它们的论文引用了7000多次。近年来DNA测序技术的快速发展促进了对单一物种种群的大规模测序。现在有关于20多万个人类个体、人类癌症、细菌菌株、主要粮食作物(例如小麦和大麦)和动物(例如奶牛)的变异的公开数据。虽然到目前为止,大多数努力都集中在利用这些数据来识别与遗传疾病有关的变异,但变异数据提供了一种全新的资源来了解蛋白质结构、功能和物种内相互作用的细节。我们小组最近的工作(MacGowan等人,2017)已经证明,变异数据可以在200多个蛋白质结构域家族中识别在蛋白质-配体特异性和蛋白质-蛋白质相互作用特异性中重要的关键残基。这个博士项目将建立在这些发现的基础上,通过包括分子动力学模拟在内的各种技术来描述所识别的位置,以确定哪些最有可能影响分子功能。该项目最初将重点放在重复家庭上,因为这些家庭为特定部位提供了最具统计学意义的信号,也是我们的联合主管Ulrich Zachariae博士的研究重点,他也是MD模拟技术方面的专家。这个项目将培训学生软件开发和先进的生物信息学研究技术,包括机器学习、NoSQL技术和统计学。完成博士学位后,学生将为生物信息学的研究生涯做好准备,同时也拥有与大数据分析或软件工程职业相适应的出色的可转移技能。
英文摘要
Our research has focused for more than 20 years on developing effective computational methods to predict the function, structure and specificity of proteins from the amino acid sequence. This experience is encapsulated in widely used BBSRC funded software tools which include the Jalview (www.jalview.org) sequence analysis workbench that has over 70,000 regular users world-wide and JPred (www.compbio.dundee.ac.uk/jpred) which performs up to 250,000 predictions/month of secondary structure and other features from the amino acid sequence for scientists in laboratories in the UK and internationally. Together, Jalview and JPred have accumulated over 7,000 citations to the papers that describe them. The rapid advances in DNA sequencing technology over recent years have stimulated the large-scale sequencing of populations of single species. There is now publicly available data on variation in over 200,000 human individuals, human cancers, bacterial strains, major food crops (e.g. wheat and barley) and animals (e.g. cow). While most effort to date has focussed on exploiting these data to identify variants involved in genetic disease, the variation data provides a completely new resource to inform details of protein structure, function and interactions within a species. Recent work from our group (MacGowan et al, 2017) has demonstrated that variation data can identify key residues important in protein-ligand specificity and protein-protein interaction specificity in over 200 protein domain families. This Ph.D. project will build on these findings to characterise the identified sites by a variety of techniques including molecular dynamics simulations to identify which are most likely to affect molecular function. The project will focus initially on repeat families since these provide the most statistically significant signals for specificity sites and are a focus of the research of our co-supervisor Dr Ulrich Zachariae who is also an expert in MD simulation techniques. This project will train the student in software development and advanced bioinformatics research techniques including machine learning noSQL technology and statistics. On completion of the Ph.D. the student will be well prepared for a research career in bioinformatics, but also have excellent transferrable skills appropriate to careers in Big Data analytics or software engineering.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
背根神经节中Mrgprd通过一种特异性lncRNA调控阿片类药物耐受的外周机制研究
-
批准号:82371224
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:马柯
-
依托单位:
多盘科单殖吸虫宿主特异性及其与无尾两栖类宿主协同进化关系研究
-
批准号:30960049
-
项目类别:地区科学基金项目
-
资助金额:23.0万元
-
批准年份:2009
-
负责人:范丽仙
-
依托单位:
Dyrk1A调控CaMKⅡδ的可变剪接及其在心脏重构过程中的作用
-
批准号:30971223
-
项目类别:面上项目
-
资助金额:31.0万元
-
批准年份:2009
-
负责人:朱健华
-
依托单位: