Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
批准号:
MR/S003711/2
负责人:
Deepti Gurdasani
金额:
$22.63万
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
大规模电子健康记录(EHR)数据有可能改变我们对疾病病因学和临床风险的理解,为临床决策和人口规模的卫生政策提供重要信息。整合多种来源的数据,包括遗传、表观遗传和临床数据,可以大大提高我们对疾病风险的理解。这样的大数据可以为精准医疗算法的开发以及与疾病相关的候选基因的新发现提供独特的机会;为基因发现与医学临床应用的整合提供一个框架。虽然大规模电子病历资源已被用于风险预测和遗传关联的识别,但对这些数据的分析仍然有限,并且没有利用数据的多维和纵向丰富性。最近在英国大规模生物数据资源的发展需要灵活的分析方法的并行发展,可以实现这些丰富的数据集的全部潜力。机器学习方法提供了一个框架,通过识别和利用最能预测结果的特征,从数据本身学习不同变量和数据类型之间的关系。这种方法可能比传统的统计方法更有优势,在传统的统计方法中,变量之间的关系需要预先说明,而且一次只能对有限数量的因素进行建模。相比之下,机器学习方法不仅允许无模型预测临床风险,而且还有助于更好地了解在大量潜在预测因素中哪些因素会影响临床风险。该提案侧重于统计和机器学习方法的开发和验证,这些方法利用复杂的数据来灵活地预测多个疾病领域患者预后的风险。此外,这些方法的应用还可以更好地了解与疾病相关的临床和遗传风险因素。这项工作将特别侧重于系统地评估和扩展当前用于风险预测的统计方法,以结合机器学习方法,这些方法可以灵活地使用大规模遗传和临床数据来模拟复杂的风险模式,同时适当地考虑风险因素的时间依赖性背景(例如重复测量)。该项目的第一阶段将涉及整合、管理和协调公共可用的生物数据资源,如英国生物银行、英国基因组学和INTERVAL研究。进一步的数据,包括基因表达和功能数据将被分层,以开发一个丰富的多维数据集。这些数据可以从序列数据中合理地预测,当不直接测量时,使用已发表的imputation和深度学习方法。在先前电子病历工作的基础上,复杂的统计方法将用于开发特定疾病领域的预测算法。下一阶段将涉及评估现有的机器和深度学习方法,这些方法可以模拟风险。然后将这些方法扩展到纵向数据模型,其中包括随时间的重复测量。这些方法的预测准确性将使用独立的数据集进行评估。为了了解疾病的遗传病因,将经典的GWAS方法与集成机器学习的方法进行比较,以允许对最重要的遗传和临床风险预测因子进行优先排序。该项目将为大数据分析背景下的临床风险预测和疾病遗传关联识别提供一个广泛的分析框架。从长远来看,这将有助于发展研究能力和专业知识的方案,以高通量的多维数据分析为目标,支持临床决策和改善患者健康。
英文摘要
Large-scale electronic health record (EHR) data has the potential to transform our understanding of disease aetiology and clinical risk, providing important information for clinical decision making and health policy at population scale. Aligning multiple sources of data, including genetic, epigenetic and clinical data can substantially improve our understanding of disease risk. Such big data can provide unique opportunities for development of algorithms for precision medicine as well as for novel discovery of candidate genes associated with disease; providing a framework for integration of genetic discovery into clinical applications in medicine. While large-scale EHR resources have been utilised for risk prediction, and identification of genetic associations, analyses of these data have been limited and have not harnessed the multi-dimensional and longitudinal richness of data. The recent development of large-scale biodata resources within the UK necessitates the parallel development of flexible analytic methods that can realise the full potential of these rich datasets.Machine learning methods provide a framework whereby the relationships among different variables and types of data are learnt from the data itself by identifying and utilising features that best predict outcomes. Such approaches may provide an advantage over classical statistical methods, where the relationships among variables require pre-specification, and only a limited number of factors can be modelled at a time. By contrast, machine learning methods not only allow model-free prediction of clinical risk, but also help better understand which factors, among large numbers of potential predictors, influence clinical risk.This proposal focuses on development and validation of statistical, and machine learning approaches that utilise complex data to flexibly predict the risk of patient outcomes across multiple disease areas. Additionally, application of such methods also allows a better understanding of clinical and genetic risk factors associated with disease. This work will specifically focus on systematically evaluating and extending current statistical methods for risk prediction to incorporate machine learning approaches that can model complex patterns of risk using large-scale genetic and clinical data flexibly, while appropriately accounting for the time-dependent context of risk factors (e.g. repeated measurements).The first phase of the project will involve integration, curation and harmonisation of publicly available biodata resources, such as the UK Biobank, Genomics England and INTERVAL study. Further data, including gene expression, and functional data will be layered to develop a rich multi-dimensional dataset. These data can be reasonably predicted from sequence data, when not directly measured, using published imputation and deep learning approaches. Developing on previous work with EHRs, complex statistical approaches will be used to develop predictive algorithms for specific disease areas. The next stage will involve evaluation of existing machine and deep learning approaches that can model risk. These approaches will then be extended to model longitudinal data incorporating repeated measurements over time. The predictive accuracy of these approaches will be evaluated using independent datasets. To understand the genetic aetiology of disease, classical GWAS approaches will be compared with approaches that integrate machine learning to allow prioritisation of the most important genetic and clinical predictors of risk.This project will provide a broad analytic framework for clinical risk prediction and identification of genetic associations with disease in the context of big data analytics. In the longer term, this will contribute to a programme of developing research capacity and expertise in high throughput analytics of multi-dimensional data with the aim of supporting clinical decision making, and improving patient health.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21203/rs.3.rs-53594/v1
发表时间:
2020
期刊:
影响因子:
--
作者:
[Alaa A]
通讯作者:
Alaa A
Assessment of COVID-19 as the Underlying Cause of Death Among Children and Young People Aged 0 to 19 Years in the US.
对 COVID-19 作为美国 0 至 19 岁儿童和青少年根本死因的评估。
DOI:
10.1001/jamanetworkopen.2022.53590
发表时间:
2023-01-03
期刊:
JAMA NETWORK OPEN
影响因子:
13.8
作者:
[Flaxman, Seth, Whittaker, Charles, Semenova, Elizaveta, Rashid, Theo, Parks, Robbie M., Blenkinsop, Alexandra, Unwin, H. Juliette T., Mishra, Swapnil, Bhatt, Samir, Gurdasani, Deepti, Ratmann, Oliver]
通讯作者:
Ratmann, Oliver
Uncertainty around the Long-Term Implications of COVID-19.
Covid-19的长期影响的不确定性。
DOI:
10.3390/pathogens10101267
发表时间:
2021-10-01
期刊:
Pathogens (Basel, Switzerland)
影响因子:
--
作者:
[Desforges M, Gurdasani D, Hamdy A, Leonardi AJ]
通讯作者:
Leonardi AJ
DOI:
10.1016/s0140-6736(20)32642-8
发表时间:
2021-01-02
期刊:
Lancet (London, England)
影响因子:
--
作者:
[Burgess RA, Osborne RH, Yongabi KA, Greenhalgh T, Gurdasani D, Kang G, Falade AG, Odone A, Busse R, Martin-Moreno JM, Reicher S, McKee M]
通讯作者:
McKee M
Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
-
批准号:MR/S003711/1
-
项目类别:Fellowship
-
资助金额:$40.71万
-
财政年份:2018
-
负责人:Deepti Gurdasani
-
依托单位:
海外基金