A Novel Reduction-Based Approach to Machine Learning Survival Modelling
A Novel Reduction-Based Approach to Machine Learning Survival Modelling
批准号:
2064211
负责人:
金额:
$0.0万
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
我的研究基础是时间序列和事件发生时间(生存)分析。其动机在于将最先进的机器学习模型引入生存分析。已经发表了几篇关于这个主题的论文,模型以前已经被用来在时间序列中使用机器学习。例如,使用高斯过程进行生存分析(van der Schaar等人,2017),或随机生存森林(Ishwaran等人,2008),仅举几个例子。然而,还没有一个全面的框架,允许在生存分析中进行严格的模型选择、验证和比较。继续之前的研究,我的博士学位将建立一个架构,从数学上并使用包括R和Python在内的软件,以允许全面的机器学习方法来进行生存分析。这将包括讨论元策略,如模型调整和集合方法。我还将讨论并尝试解决时间序列数据中出现的问题,如班级不平衡和在线更新。当我们研究需要大量时间进行培训的模型时,在线建模的重要性尤其重要。通过利用更新,我们将尝试消除重新培训过程并提高模型效率。我的研究将有两个主要目标,大致可分为元分析和模型创建。首先,我将对常用的生存模型进行全面的研究,如考克斯比例风险模型,并对照利用机器学习的更现代的模型对这些模型进行评估。这将包括研究用于评估不同类型的生存模型的指标之间的关系。此外,我将研究解决时间序列中出现的问题的常用技术,如审查和不平衡。第二部分的研究将建立在第一部分的基础上,以及我之前对减少机器学习生存任务的研究。在这里,我将推出一个全面的工作流程,用于模型的选择、评估和比较。这项研究是相关的,因为有许多问题仍然没有得到回答。例如,虽然元策略(如调整)在更经典的监督学习模型中得到了很好的研究和理解,但在调整生存模型方面的研究较少。此外,常用的生存模型经常相互比较以评估性能,并可以计算各种残差统计数据,但对于生存模型来说,还没有一个明确的指标来表示一个好的性能统计数据应该是什么样子。例如,一个进行随机概率预测的分类模型的Brier得分将在0.25左右,因此任何低于这一得分的模型都可以被认为是“好的”,但你能赋予偏差值为40,000的单个Cox模型什么意义?这个缺失的框架对于使机器学习中的生存建模达到与更经典的监督设置相同的标准至关重要。由于我最初的重点将是基于约简的方法,我还将研究一些开放的问题,例如在生存建模的背景下定义约简意味着什么,以及这些模型如何在数学上与监督方法相关(例如,考克斯模型和广义线性模型之间的联系已经很好地理解了)。生存分析在患者保健数据和预测患者未来健康状态(疾病风险、中风、死亡等)的背景下是最重要的。考虑到这一点,我将根据真实世界的医疗保健数据以及可以在各种情况下测试模型的合成数据来评估所有新的和现有的模型。这可能会以大数据的形式增加一层困难,因为横截面医疗数据集,特别是那些具有多个时间点的数据集,可能很快就会变得非常大;我希望在我的研究中也解决这一问题。
英文摘要
The foundation of my research rests in time-series and time-to-event (Survival) analysis. The motivation lies in bringing state-of-the-art machine learning models to Survival analysis. Several papers have been published on this subject and models have been previously utilised to use machine learning in time-series. For example, the use of Gaussian Processes for survival analysis (van der Schaar et al., 2017), or Random Survival Forests (Ishwaran et al., 2008), just to name a couple. However there has yet to be a comprehensive framework that allows for rigorous model selection, validation and comparison in Survival analysis.Continuing previous research my PhD will be building an architecture, both mathematically and using software including R and Python, to allow for a comprehensive machine learning approach to Survival Analysis. This will include discussing meta-strategies such as model tuning and ensemble-methods. I will also discuss and attempt to solve problems that arise with time-series data, such as class imbalance and online updating. The importance of online modelling is particularly relevant when we look at models that can take an extensive period of time for training. By utilising updating we will attempt to remove the re-training process and improve model efficiency.My research will have two primary aims that can be roughly split into meta-analysis and model creation. In the first instance I will create a comprehensive study of commonly used Survival models, such as Cox Proportional Hazards, and assess these against more modern models that make use of machine learning. This will include studying the relationships between the metrics used to evaluate differing types of Survival Models. Additionally, I will be looking at commonly used techniques to solve problems that arise in time-series, such as censoring and imbalance. The second part of the research will build on the first part as well as my previous research into reduction of the machine learning Survival task. Here I will derive a comprehensive workflow for model selection, evaluation and comparison. This research is relevant as there are many questions that remain unanswered. For example, whilst meta-strategies such as tuning are well-researched and understood in more classical supervised learning models, less research has been placed in tuning Survival models. Moreover, commonly used Survival models are often compared to each other to assess performance and various residual statistics can be computed but there is yet to be a well-defined indication of what a 'good' performance statistic for a Survival model would look like. For example, a classification model that makes random probabilistic predictions will achieve a Brier score around 0.25, so any model below this can be considered 'good', but what meaning can you give to a single Cox model with a Deviance of 40,000? This missing framework is vital to bringing Survival modelling in machine learning up to the same standards as the more classical supervised setting. As my initial focus will be on a reduction-based approach I will also be looking at open questions such as defining what reduction means in the context of Survival modelling and how these models can be mathematically related to supervised approaches (for example the connection between the Cox model and generalized linear models are already well understood). Survival analysis is most important in the context of patient health-care data and predicting the future health-state of a patient (risk of illness, stroke, death, etc.). With this in mind I will evaluate all new and existing models against both real-world health-care data as well as synthesised data that can test the models in a wide variety of cases. This will likely add a layer of difficulty in the form of Big Data as cross-sectional healthcare data-sets, especially as those with multiple time-points, can quickly become very large; I hope to also tackle this in my research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
兼捕减少装置(Bycatch Reduction Devices, BRD)对拖网网囊系统水动力及渔获性能的调控机制
-
批准号:32373187
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:唐浩
-
依托单位: