Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
批准号:
MR/S003711/1
负责人:
Deepti Gurdasani
金额:
$40.71万
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
大规模电子健康记录(EHR)数据有可能改变我们对疾病病因和临床风险的理解,为临床决策和人群规模的卫生政策提供重要信息。整合多种数据来源,包括遗传、表观遗传和临床数据,可以大大提高我们对疾病风险的理解。这样的大数据可以为精确医学算法的开发以及与疾病相关的候选基因的新发现提供独特的机会;为将遗传发现整合到医学临床应用中提供框架。虽然大规模的EHR资源已被用于风险预测和遗传关联的识别,但对这些数据的分析有限,并且没有利用数据的多维和纵向丰富性。最近英国大规模生物数据资源的发展需要灵活的分析方法的并行开发,可以实现这些丰富的datasets.Machine learning方法提供了一个框架,通过识别和利用最佳预测结果的功能,从数据本身学习不同变量和数据类型之间的关系。这种方法可以提供优于经典统计方法的优势,在经典统计方法中,变量之间的关系需要预先指定,并且一次只能模拟有限数量的因素。相比之下,机器学习方法不仅可以无模型地预测临床风险,还可以帮助更好地了解在大量潜在的预测因素中,哪些因素会影响临床风险。该提案侧重于开发和验证统计和机器学习方法,这些方法利用复杂的数据灵活地预测多个疾病领域的患者结局风险。此外,这些方法的应用还允许更好地理解与疾病相关的临床和遗传风险因素。这项工作将特别侧重于系统地评估和扩展目前用于风险预测的统计方法,以纳入机器学习方法,这些方法可以灵活地使用大规模遗传和临床数据对复杂的风险模式进行建模,同时适当考虑风险因素的时间依赖性背景。(例如重复测量)。该项目的第一阶段将涉及整合、管理和协调公开可用的生物数据资源,如英国生物库,Genomics England和INTERVAL研究。进一步的数据,包括基因表达和功能数据将分层,以开发丰富的多维数据集。当没有直接测量时,可以使用已发表的插补和深度学习方法从序列数据合理预测这些数据。在以前EHR工作的基础上,复杂的统计方法将用于开发特定疾病领域的预测算法。下一阶段将涉及评估现有的机器和深度学习方法,这些方法可以对风险进行建模。然后,这些方法将扩展到模型纵向数据,包括随着时间的推移重复测量。这些方法的预测准确性将使用独立的数据集进行评估。为了了解疾病的遗传病因学,将比较经典的GWAS方法与集成机器学习的方法,以优先考虑最重要的遗传和临床风险预测因子。该项目将在大数据分析的背景下为临床风险预测和识别与疾病的遗传关联提供广泛的分析框架。从长远来看,这将有助于发展多维数据高通量分析的研究能力和专业知识,旨在支持临床决策和改善患者健康。
英文摘要
Large-scale electronic health record (EHR) data has the potential to transform our understanding of disease aetiology and clinical risk, providing important information for clinical decision making and health policy at population scale. Aligning multiple sources of data, including genetic, epigenetic and clinical data can substantially improve our understanding of disease risk. Such big data can provide unique opportunities for development of algorithms for precision medicine as well as for novel discovery of candidate genes associated with disease; providing a framework for integration of genetic discovery into clinical applications in medicine. While large-scale EHR resources have been utilised for risk prediction, and identification of genetic associations, analyses of these data have been limited and have not harnessed the multi-dimensional and longitudinal richness of data. The recent development of large-scale biodata resources within the UK necessitates the parallel development of flexible analytic methods that can realise the full potential of these rich datasets.Machine learning methods provide a framework whereby the relationships among different variables and types of data are learnt from the data itself by identifying and utilising features that best predict outcomes. Such approaches may provide an advantage over classical statistical methods, where the relationships among variables require pre-specification, and only a limited number of factors can be modelled at a time. By contrast, machine learning methods not only allow model-free prediction of clinical risk, but also help better understand which factors, among large numbers of potential predictors, influence clinical risk.This proposal focuses on development and validation of statistical, and machine learning approaches that utilise complex data to flexibly predict the risk of patient outcomes across multiple disease areas. Additionally, application of such methods also allows a better understanding of clinical and genetic risk factors associated with disease. This work will specifically focus on systematically evaluating and extending current statistical methods for risk prediction to incorporate machine learning approaches that can model complex patterns of risk using large-scale genetic and clinical data flexibly, while appropriately accounting for the time-dependent context of risk factors (e.g. repeated measurements).The first phase of the project will involve integration, curation and harmonisation of publicly available biodata resources, such as the UK Biobank, Genomics England and INTERVAL study. Further data, including gene expression, and functional data will be layered to develop a rich multi-dimensional dataset. These data can be reasonably predicted from sequence data, when not directly measured, using published imputation and deep learning approaches. Developing on previous work with EHRs, complex statistical approaches will be used to develop predictive algorithms for specific disease areas. The next stage will involve evaluation of existing machine and deep learning approaches that can model risk. These approaches will then be extended to model longitudinal data incorporating repeated measurements over time. The predictive accuracy of these approaches will be evaluated using independent datasets. To understand the genetic aetiology of disease, classical GWAS approaches will be compared with approaches that integrate machine learning to allow prioritisation of the most important genetic and clinical predictors of risk.This project will provide a broad analytic framework for clinical risk prediction and identification of genetic associations with disease in the context of big data analytics. In the longer term, this will contribute to a programme of developing research capacity and expertise in high throughput analytics of multi-dimensional data with the aim of supporting clinical decision making, and improving patient health.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21203/rs.3.rs-53594/v1
发表时间:
2020
期刊:
影响因子:
--
作者:
[Alaa A]
通讯作者:
Alaa A
The government wants us to learn to live with covid-19, but where is the learning?
政府希望我们学会如何与 covid-19 共存,但学习在哪里?
DOI:
10.1136/bmj.o1096
发表时间:
2022
期刊:
BMJ (Clinical research ed.)
影响因子:
--
作者:
[Gurdasani D]
通讯作者:
Gurdasani D
Uncertainty around the Long-Term Implications of COVID-19.
Covid-19的长期影响的不确定性。
DOI:
10.3390/pathogens10101267
发表时间:
2021-10-01
期刊:
Pathogens (Basel, Switzerland)
影响因子:
--
作者:
[Desforges M, Gurdasani D, Hamdy A, Leonardi AJ]
通讯作者:
Leonardi AJ
Assessment of COVID-19 as the Underlying Cause of Death Among Children and Young People Aged 0 to 19 Years in the US.
对 COVID-19 作为美国 0 至 19 岁儿童和青少年根本死因的评估。
DOI:
10.1001/jamanetworkopen.2022.53590
发表时间:
2023-01-03
期刊:
JAMA NETWORK OPEN
影响因子:
13.8
作者:
[Flaxman, Seth, Whittaker, Charles, Semenova, Elizaveta, Rashid, Theo, Parks, Robbie M., Blenkinsop, Alexandra, Unwin, H. Juliette T., Mishra, Swapnil, Bhatt, Samir, Gurdasani, Deepti, Ratmann, Oliver]
通讯作者:
Ratmann, Oliver
DOI:
10.1016/s0140-6736(20)32642-8
发表时间:
2021-01-02
期刊:
Lancet (London, England)
影响因子:
--
作者:
[Burgess RA, Osborne RH, Yongabi KA, Greenhalgh T, Gurdasani D, Kang G, Falade AG, Odone A, Busse R, Martin-Moreno JM, Reicher S, McKee M]
通讯作者:
McKee M
Predictive analytics of integrated genomic and clinical data using machine learning and complex statistical approaches
-
批准号:MR/S003711/2
-
项目类别:Fellowship
-
资助金额:$22.63万
-
财政年份:2019
-
负责人:Deepti Gurdasani
-
依托单位:
海外基金