Scalable Bayesian Statistical Machine Learning for Multi-modal Data with Applications to Multiple Sclerosis
Scalable Bayesian Statistical Machine Learning for Multi-modal Data with Applications to Multiple Sclerosis
批准号:
2740724
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
在这个项目中,我们承担了有效管理广泛和复杂的数据集的挑战,NO.MS临床数据集就是一个例子。主要目标是通过揭示潜在变量来简化这些数据集的内在复杂性。这种简化在处理高维数据时尤其重要,因为高维数据的挑战在于将丰富的维度提炼成一组更简洁的更广泛的协变量,这些变量仍然是数据的准确表示。这个提取潜在变量的概念将其相关性扩展到我们直接研究的不同领域。由诺华公司提供的NO.MS数据集是一个感兴趣的焦点。它包含了大量关于多发性硬化症患者的数据,使其成为同类数据中最大的数据集之一。这种区别源于包含了大量的MRI脑部扫描,由于每次扫描中有无数像素,这导致了它的高维度。因此,这个数据集具有很大的潜力来揭开对这种疾病的洞察。然而,从分析的角度来看,它提出了相当大的挑战。这种复杂性源于数据集将离散数据(如残疾评分)和连续数据(例如MRI扫描中的像素值)合并在一起。此外,该数据集来自多项研究,每项研究都捕获了患者就诊的不同方面,导致大量数据缺失。该项目的核心目标包括两个关键贡献:首先,我们的目标是构建一个模型,能够有效地揭示低维潜在空间的可解释表示。我们的方法在很大程度上依赖于贝叶斯统计,这是一个将先前的信念整合到建模中的统计框架,随后根据传入的数据对其进行更新。该模型必须具有适应连续和离散数据的通用性,为像NO.MS这样的数据集提供解决方案。此外,我们优先考虑可伸缩性,认识到传统方法管理大型高维数据集的不切实际,如诺华多发性硬化症数据。此外,我们的模型应该自动确定准确表示数据所需的最佳潜在变量数量。虽然现有的模型可能会解决这些挑战的个别方面,但我们方法的独特方面是将解决方案集成到一个具有凝聚力的整体中。第二,我们打算将这个全面的模型应用于NO.MS数据集,以加深我们对多发性硬化症的理解。这可以通过与医学专家合作分析已揭示的潜在因素来实现。另外。这些潜在因素提供了数据的简化表示,进而可以与计算更密集的模型结合使用。与传统方法相比,这种简化的表示法提高了我们分析的效率。该项目属于EPSRC统计和应用概率研究领域,并与诺华公司合作进行,由Habib Ganjgahi博士、Tom Nichols教授和Chris Holmes教授监督。
英文摘要
Within this project, we undertake the challenge of effectively managing extensive and complex datasets, as exemplified by the NO.MS clinical dataset. The primary objective is to streamline the inherent complexity of these datasets by uncovering underlying latent variables. This simplification is particularly vital when contending with high-dimensional data, where the challenge lies in distilling an abundance of dimensions into a more concise set of broader covariates that remain accurate representations of the data. This concept of distilling latent variables extends its relevance to diverse fields beyond our immediate study.The NO.MS dataset, provided by Novartis, serves as a focal point of interest. It encompasses a wealth of data on individuals affected by multiple sclerosis, distinguishing itself as one of the largest datasets of its kind. This distinction arises from the inclusion of numerous MRI brain scans, contributing to its high dimensionality due to the myriad of pixels within each scan. Consequently, this dataset bears significant potential for unravelling insights into the disease. However, from an analytical standpoint, it presents considerable challenges. This complexity stems from the dataset's amalgamation of discrete data, such as disability scores, and continuous data, exemplified by the pixel values within the MRI scans. Furthermore, the dataset draws from multiple studies, each capturing distinct facets of patient visits, resulting in a substantial amount of missing data.The project's core objectives comprise of two pivotal contributions:First and foremost, we aim to construct a model that can effectively unveil an interpretable representation of the lower-dimensional latent space. Our approach relies heavily on Bayesian statistics, a statistical framework that integrates prior beliefs into modelling, subsequently updating them based on the incoming data. This model must possess the versatility to accommodate both continuous and discrete data, offering a solution for datasets like NO.MS. Furthermore, we prioritize scalability, recognizing the impracticality of conventional methods for managing large, high-dimensional datasets, such as the Novartis multiple sclerosis data. In addition, our model should autonomously determine the optimal number of latent variables required to represent the data accurately. While existing models may address individual aspects of these challenges, the unique aspect of our approach is the integration of solutions into a cohesive whole.Secondly, we intend to apply this comprehensive model to the NO.MS dataset to deepen our understanding of multiple sclerosis. This can be achieved by analysing the latent factors unveiled, in collaboration with medical experts. Additionally. These latent factors furnish a simplified representation of the data, which can, in turn, be employed in conjunction with more computationally intensive models. This streamlined representation enhances the efficiency of our analyses compared to conventional approaches.This project falls within the EPSRC Statistics and Applied Probability research area and is carried out in collaboration with Novartis, it is supervised by Dr Habib Ganjgahi, Prof Tom Nichols and Prof Chris Holmes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
基于 Bayesian 动态权重的脑出血早期风险预测模型方法研究
-
批准号:JCZRQNB202600722
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
多元纵向数据与复发事件和终止事件的Bayesian联合模型研究
-
批准号:82173628
-
项目类别:面上项目
-
资助金额:52万元
-
批准年份:2021
-
负责人:尹平
-
依托单位:
三维地质模型约束下地球化学场的Bayesian-MCMC推断
-
批准号:42072326
-
项目类别:面上项目
-
资助金额:63.0万元
-
批准年份:2020
-
负责人:张宝一
-
依托单位:
基于Bayesian Kriging模型的压射机构稳健优化设计基础研究
-
批准号:51875209
-
项目类别:面上项目
-
资助金额:59.0万元
-
批准年份:2018
-
负责人:游东东
-
依托单位:
X射线图像分析中的MCMC-Bayesian理论与计算方法研究
-
批准号:U1830105
-
项目类别:联合基金项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:李庆武
-
依托单位:
基于Bayesian位移场的SAR图像精确配准方法研究
-
批准号:41601345
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2016
-
负责人:丁明涛
-
依托单位:
多结局Bayesian联合生存模型及糖尿病并发症预测研究
-
批准号:81673274
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2016
-
负责人:余小金
-
依托单位:
基于Meta流行病学和Bayesian方法构建针刺干预无偏倚风险效果评价体系研究
-
批准号:81403276
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2014
-
负责人:杜亮
-
依托单位:
BtoC电子商务中基于分层Bayesian网络的信任与声誉计算理论研究
-
批准号:71302080
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2013
-
负责人:田博
-
依托单位:
基于Bayesian网络的坚硬顶板条件下煤与瓦斯突出预警控制机理研究
-
批准号:51274089
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:杨玉中
-
依托单位: