Computational analysis of protein covariation for the identification of disease-associated variants in coding regions
Computational analysis of protein covariation for the identification of disease-associated variants in coding regions
批准号:
MR/R010900/1
负责人:
David Talavera
金额:
$48.93万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
为每个人实现良好的健康和幸福是非常多样化的倡议的主要目标。当地和全球机构的目标是在今后几年降低死亡率和提高护理标准。实现这些目标的一个关键因素是开发更好的诊断工具;也就是说,更准确的诊断可能意味着更精确的治疗,更少的次要影响。由于大多数人类疾病都有遗传因素,因此我们必须改进计算方法,以准确识别导致先天性疾病(如囊性纤维化)或增加患多因素疾病(如2型糖尿病)的风险的基因组变异。目前发现致病变异和遗传诊断的方法主要集中在识别基因组变异:1)在患者中存在,但在健康人群中不存在或极其罕见,以及2)基于它们缺乏进化保守而被认为是有害的。这两种情况都是基于单一变异导致这种疾病的假设。然而,在许多情况下,关于导致疾病的特定变异,对数据的分析仍然没有定论。我的假设是,在许多情况下,这种疾病不是由单一的变异引起的,而是由不利的基因组变异的组合引起的;也就是说,在人类群体中单独发现的变异,但当在一个人身上发现时,会造成有害的影响。我将重点介绍编码基因组(基因组中编码蛋白质的部分)中的变异,因为蛋白质执行细胞和组织中的绝大多数生物过程。首先,我将确定哪些蛋白质位置显示出在整个进化过程中或在人类群体中存在协同变异的证据;即那些位置对,通过这些位置,为了保持有机体的适应性,经常同时发生变化。其次,我将确定哪些氨基酸组合在这些不同的位置上更受青睐。第三,我将使用这些信息重新分析NHS患有心脏或眼睛遗传疾病的患者的基因测试数据。我的目标是提高可以通过基因诊断的病例比率。最后,我将分析法洛四联症患者的测序数据,这是一种先天性心脏病的典型。这是一种复杂的遗传疾病,涉及各种基因的变异;然而,确切的原因尚不清楚。我将确定哪些病例是由有害的变种组合引起的。我的研究将产生一系列计算工具,将免费提供给学术界和临床界。这些工具将为实现更好的基因诊断目标做出决定性的贡献。
英文摘要
Attaining Good Health and Well-Being for everyone are the main goals of very diverse initiatives. Local and global Institutions are targeting a reduction in mortality and an improvement of standard of care for the forthcoming years. A key element in achieving these goals is the development of better diagnostics tools; i.e., moreaccurate diagnoses may mean more precise treatments with fewer secondary effects.Because most of the human diseases have a genetic factor, it is essential that we improve the computational methods for precisely identifying the genome variants that cause congenital disorders (e.g. cystic fibrosis), or increase the risk of suffering multifactorial diseases such as type-2 diabetes.Current approaches for discovery of disease-causing variants and for genetic diagnostics focus in the identification of genome variants that 1) are present in patients, but non-existing or extremely rare in the healthy populations, and 2) are deemed detrimental based on their lack of evolutionary conservation. Both of these conditions are implicitly based on the assumption that single variants are responsible for the disease. Nevertheless, in many cases the analyses of data remain inconclusive in regards of the particular variants that cause the disorders.My hypothesis is that in many cases the disease is not caused by single variants, but by an unfavourable combination of genomic variants; that is, variants that are separately found in the human population, but that cause a harmful effect when found together in an individual. I will focus on variants within the coding genome (thepart of the genome that codifies for proteins), because proteins carry out the vast majority of biological processes within cells and tissues. First, I am going to identify which protein positions show evidence of covariation throughout evolution or withinthe human population; namely, those pairs of positions whereby changes have often occurred concurrently in order to maintain the fitness of the organism. Second, I am going to identify which combinations of amino acids are favoured within those covarying positions. Third, I am going to use this information for reanalysing genetic testing data from NHS patients suffering of cardiac or eye genetic disorders. My goal is to increase the rate of cases that can be genetically diagnosed. Finally, I will analyse sequencing data from patients suffering of Tetralogy of Fallot, which is atype of Congenital Heart Disease. This is a complex genetic disorder involving variants in various genes; however, the exact causes are unknown. I will identify cases where the disorder is caused by a harmful combination of variants. My research is going to produce a series of computational tools that will be made freely available to the academic and clinical communities. These tools will make adecisive contribution to the goal of achieving better genetic diagnoses.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1038/s10038-022-01051-y
发表时间:
2022-10
期刊:
Journal of human genetics
影响因子:
3.5
作者:
[Chelu A, Williams SG, Keavney BD, Talavera D]
通讯作者:
Talavera D
DOI:
10.1038/s41598-022-21433-8
发表时间:
2022-11-04
期刊:
SCIENTIFIC REPORTS
影响因子:
4.6
作者:
[Byrne, Dominic J. F., Williams, Simon G., Nakev, Apostol, Frain, Simon, Baross, Stephanie L., Vestbo, Jorgen, Keavney, Bernard D., Talavera, David]
通讯作者:
Talavera, David
Modelling Human Genomic Variations Using Markov Random Field: A Feasibility Study
-
批准号:EP/W016109/1
-
项目类别:Research Grant
-
资助金额:$10.25万
-
财政年份:2022
-
负责人:David Talavera
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:USHARANI HAREESH GOVINDARA JAN
-
依托单位:
利用全基因组关联分析和QTL-seq发掘花生白绢病抗性分子标记
-
批准号:31971981
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:晏立英
-
依托单位:
基于SERS纳米标签和光子晶体的单细胞Western Blot定量分析技术研究
-
批准号:31900571
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:刘兵
-
依托单位:
利用多个实验群体解析猪保幼带形成及其自然消褪的遗传机制
-
批准号:31972542
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2019
-
负责人:郭源梅
-
依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
-
批准号:41601604
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:赵爱琴
-
依托单位:
基于个体分析的投影式非线性非负张量分解在高维非结构化数据模式分析中的研究
-
批准号:61502059
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2015
-
负责人:刘昶
-
依托单位:
多目标诉求下我国交通节能减排市场导向的政策组合选择研究
-
批准号:71473155
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2014
-
负责人:柴建
-
依托单位:
大规模微阵列数据组的meta-analysis方法研究
-
批准号:31100958
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:赵洪雅
-
依托单位:
基于物质流分析的中国石油资源流动过程及碳效应研究
-
批准号:41101116
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2011
-
负责人:刘晓洁
-
依托单位: