Artificial Intelligence for Missing Data Imputation in Electronic Medical Records
Artificial Intelligence for Missing Data Imputation in Electronic Medical Records
批准号:
NE/T013982/1
负责人:
Alastair Denniston
金额:
$1.31万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --
中文摘要
英国和加拿大的医疗系统多年来一直广泛使用电子病历(EMR)作为其业务的一个组成部分。然而,虽然存在数字记录数据,但它们作为“学习型卫生系统”的基础,通过整理和检查数据和证据来改进所有这些领域,从而不断改善患者体验、医院操作和护理质量。然而,现实世界的EMR数据处理起来非常困难。造成这些困难的一个重要因素是数据质量。数据缺失是一个特别的问题,某些记录的缺失率在10-30%之间。正确处理EMR数据中的缺失数据问题是复杂的,因为很难区分真正的缺失数据(数据没有记录到系统中)和不适用的响应(例如,测试不合适,因此没有进行)。数据可以是随机缺失(MAR)或非随机缺失(MNAR),在后者中,存在一个决定缺失模式的潜在因素。因此,某些类型的缺失可能是“有益的”,因为如果临床医生决定不要求进行某些检查,这表明对病人的健康状况有某种隐性的信念。不考虑这些偏见的来源可能会导致不正确的推论。人工智能技术被视为解锁电子病历中信息财富的重要工具。该项目将有助于这些技术的成熟,以解决现实世界EMR数据集的复杂性。这里提出的研究将开发数据输入算法,寻求更强大,可靠和推广。我们最初选择专注于脓毒症的自动诊断,这是生物医学研究的一个紧迫领域,因为脓毒症每年仅在英国就导致约44,000人死亡。因此,通过将基于机器学习的现代方法应用于大型EMR数据集,我们承诺以一种独特的方式解决这个问题,这可能会对现实世界产生有意义的影响。然而,由于许多人工智能预测模型需要完整的数据集作为输入,处理缺失数据的一种流行策略涉及“数据输入”,即使用算法来填充缺失的数据值。这些方法的复杂性各不相同,从简单地用整个数据集的平均观测值填充缺失值,到试图引出数据中的底层模式的更高级的方法。然而,许多当前的归算方法仅针对某些类型的EMR数据(例如临床时间序列的分子测量)而设计,并且无法解释偏差的来源,也无法提供关于归算数据质量的确定性措施。该项目的总体目标是开发新的机器学习方法,用于emr中的缺失数据输入,以解释输入中的偏差和统计不确定性。
英文摘要
Health systems in the UK and Canada have made extensive use of Electronic Medical Records (EMR) for many years as an integral part of their operations. However, whilst digitally recorded data exists, their use as the basis of a "learning health system" whereby continuous improvements in patient experience, hospital operations, and quality of care has are made by collating and examining data and evidence to improve all these areas. However, real-world EMR data can be very challenging to handle.One significant contribution to these difficulties is data quality. Missing data is a particular issue, with rates of missingness of between 10-30% for some records. Properly addressing the missing data issue in EMR data is complicated by the fact that it can be difficult to differentiate between genuine missing data (data was not recorded into the system) and a non-applicable response (e.g. the test was not appropriate therefore it was not done). Data can be missing-at-random (MAR) or missing-not-at-random (MNAR) where, in the latter, there is an underlying factor that determines the missingness patterns. Certain types of missingness can therefore be "informative" since, if a clinician decided not to order certain tests, it indicates a certain implicit belief about the perceived health state of the patient. Failure to account for these sources of bias may lead to incorrect inferences.Artificial Intelligence technologies are seen as an important tool in unlocking the information wealth held in our electronic medical records. This project will contribute to the maturation of these technologies to account for the real-world complexities of EMR datasets. The research proposed here will develop algorithms for data imputation that seek to be more robust, reliable and generalisable. We have chosen to initially focus on automated sepsis diagnosis, a pressing area of biomedical research given that sepsis accounts for around 44,000 deaths each year in the UK alone. Therefore, by applying modern approaches based on machine learning to large EMR datasets we promise to tackle this problem in a unique way that could have meaningful real-world impact.However, as many AI prediction models require complete datasets as input, one popular strategy for handling missing data involves "data imputation", whereby an algorithm is used to fill in missing data values. These methods vary in complexity from simply filling in missing values with the average observed values over the entire dataset through to more advanced methods that attempt to elicit the underlying patterns in the data. However, many current imputation methods are designed for only certain types of EMR data (e.g. clinical time series of molecular measurements) and fail to account for sources of bias and provide measures of certainty about the quality of the imputed data. The overall goal of this project is to develop novel machine learning methods for missing data imputation in EMRs that account for biases and statistical uncertainty in the imputation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
INSIGHT: The Health Data Research Hub for Eye Health
-
批准号:MC_PC_19005
-
项目类别:Intramural
-
资助金额:$431.8万
-
财政年份:2019
-
负责人:Alastair Denniston
-
依托单位:
Regulation of dendritic cell function by the ocular microenvironment in uveitis
-
批准号:G0600416/1
-
项目类别:Fellowship
-
资助金额:$25.04万
-
财政年份:2006
-
负责人:Alastair Denniston
-
依托单位:
海外基金