Improving Statistical Machine Learning approaches for Time-to-Event Prediction Modelling
Improving Statistical Machine Learning approaches for Time-to-Event Prediction Modelling
批准号:
2722161
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
背景:同一个体发生一种以上慢性健康状况对长期健康结果的影响往往不清楚。现有的研究通常侧重于单一条件与单一结果的关联,往往忽略或选择可能具有多种长期条件(MLTC)背景的个体。因此,可能无法充分解决某些患有慢性肝炎的个人群体的保健需要。研究母婴保健服务影响的一个障碍是可能的健康状况组合的数量。确定和招募足够多的具有某些条件的个人进行研究可能是不可行的。为了克服这些挑战,最近的一种方法是回顾性地使用病历数据库中保存的信息。可以从数据库中确定具有相似特征的个人群体,并与其他群体进行比较,以确定为什么会出现某些健康结果。然后,可以创建基于统计机器学习算法的计算机模型,以根据个人特征预测这些健康结果的未来风险。然而,使用这些历史数据集来创建预测模型需要小心处理。数据可能是在与您可能感兴趣的情况不同的背景和/或前提下获得的,这可能导致计算机模型给出有偏见或误导性的见解。此外,复杂的机器学习模型可能缺乏鲁棒性,容易产生不稳定的行为,例如,对两个几乎相同的病人给出截然不同的风险概率。本研究旨在开发方法,以提高从观测数据构建的基于统计机器学习的预测模型的鲁棒性和有效性:1)评估现有用于时间到事件建模的统计机器学习方法的鲁棒性和稳定性,2)开发方法以改进用于时间到事件建模的统计机器学习方法的鲁棒性和稳定性,3)测试新颖的方法使用真实世界初级保健数据的方法,并将其与现有方法进行比较。研究方法的新颖性标准机器学习的发展侧重于使用与准确性相关的标准来衡量预测模型的表现。然而,越来越多的人意识到,在实际使用中,准确性只是决定预测模型有用性的几个重要标准之一。在本研究中,我们将研究模型训练标准的使用,该标准包含以下考虑因素:(i)模型稳定性的四个级别(如Riley和Collins(2023)所定义的),(ii)更新后模型版本之间的一致性,以及(iii)对异常数据输入的敏感性。该项目属于“EPSRC医疗技术研究领域”,其中“优化疾病预测、诊断和干预”是本网站列出的主题或研究领域之一https://www.epsrc.ac.uk/research/ourportfolio/themes/It将为分析大型现实世界的初级保健健康数据集创造新的方法,支持特定患者的预测模型,并支持识别预防疾病或疾病复发的机会。合作这个项目将包括与伯明翰大学的合作。
英文摘要
BackgroundThe implications of more than one chronic health condition occurring in the same individual on long-term health outcomes is often unclear. Existing studies typically focus on the association of a single condition with a single outcome and often ignore or select out individuals who might have a background of multiple long-term conditions (MLTC). Consequently, the health needs of certain groups of individuals with MLTCs may not be adequately addressed. One obstacle to studying the impact of MLTC is the number of possible combinations of health conditions. It may not be feasible to identify and recruit sufficiently large numbers of individuals with certain sets of conditions to study.To overcome these challenges, a recent approach has been to retrospectively use information held in medical record databases. Groups of individuals with similar characteristics can be identified from the database and compared with other groups to determine why certain health outcomes manifest. Computer models, based on statistical machine learning algorithms, can then be created to predict the future risk of these health outcome given individual characteristics. However, the use of this historical datasets to create prediction models needs careful handling. The data may have been acquired under a different context and/or premise to the situation in which you may be interested, and this could lead to computer models that give biased or misleading insights. In addition, complex machine learning models can lack robustness and be prone to unstable behaviour, for example, giving very different risk probabilities for two nearly identical patients.Aims & ObjectivesThis research aims to develop methodologies that will improve the robustness and validity of statistical machine learning-based prediction models that are constructed from observational data:1) To assess the robustness and stability of existing statistical machine learning approaches for time-to-event modelling,2) To develop methodology to improve aspects of the robustness and stability of statistical machine learning approaches for time-to-event modelling,3) To test the novel methodology using real-world primary care data and compare it to existing approaches.Novelty of the research methodologyStandard machine learning development focuses on the use of accuracy-related criteria to measure how well prediction models perform. However, there is increasing awareness that in real-world usage, accuracy is just one of several important criteria that determines the usefulness of a prediction model. In this research we will study the use of model training criteria that encompass considerations of the (i) four levels of model stability (as defined in Riley & Collins (2023), (ii) consistency between model versions after updating, and (iii) sensitivity to unusual data inputs.Alignment to EPSRC's strategies and research areasThis project falls within the 'EPSRC Healthcare Technologies research area' where "Optimising disease prediction, diagnosis and intervention" is one of the themes or research areas listed on this website https://www.epsrc.ac.uk/research/ourportfolio/themes/It will create new methods for analysing large real-world primary care health data sets, underpin patient-specific predictive models, and support the identification of opportunities for prevention of disease or its recurrence.CollaborationsThis project will involve a collaboration with the University of Birmingham.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金