Using Machine Learning to Identify Health Outcomes from Electronic Health Record Data

Using Machine Learning to Identify Health Outcomes from Electronic Health Record Data
复制标题

DOI:
10.1007/s40471-018-0165-9
复制
发表时间:
2018-12-01
影响因子:
3.3
通讯作者:
Toh, Sengwee
Toh, Sengwee
中科院分区:
医学4区
文献类型:
--
作者:
Wong, Jenna;Murray Horwitz, Mara;Toh, Sengwee

文献摘要

被引文献

相似文献

电子健康记录(EHR)包含用于识别健康结果的有价值的数据,但这些数据在创建可计算表型分析算法时也存在许多挑战。机器学习方法可以帮助解决其中一些挑战。在这篇综述中,我们讨论了四种常见的情况,研究人员可能会发现这些情况有助于批判性地思考机器学习何时以及在什么任务中可以用于从EHR数据中识别健康结果。最近的发现我们首先考虑机器学习在健康结果的两个维度方面可能特别有用的条件:(1)其诊断标准的特征和(2)其诊断数据通常存储在EHR系统中的格式。在第一个维度中,我们提出,对于诊断标准涉及许多临床因素、模糊定义或主观解释的健康结果,机器学习可能有助于从临床输入向量建模复杂的诊断决策过程,以识别具有健康结果的个体。在第二维度中,我们提出,对于诊断信息主要以非结构化格式(如自由文本或图像)存储的健康结果,机器学习可能有助于提取和构建这些信息,作为自然语言处理系统或图像识别任务的一部分。然后,我们将这两个维度结合起来,定义四种常见的健康结果情景。对于每种情况,我们讨论了机器学习的潜在用途-首先假设准确和完整的EHR数据,然后放松这些假设以适应现实世界EHR系统的限制。我们用具体的例子来说明这四种情况,并描述了最近的研究如何使用机器学习来识别这些健康结果从EHR data.SummaryMachine学习有很大的潜力,以提高准确性和效率的健康结果识别从EHR系统,特别是在某些条件下。为了促进机器学习在基于EHR的表型分析任务中的使用,未来的工作应该优先考虑提高机器学习算法在多站点环境中使用的可移植性。
Purpose of ReviewElectronic health records (EHRs) contain valuable data for identifying health outcomes, but these data also present numerous challenges when creating computable phenotyping algorithms. Machine learning methods could help with some of these challenges. In this review, we discuss four common scenarios that researchers may find helpful for thinking critically about when and for what tasks machine learning may be used to identify health outcomes from EHR data.Recent FindingsWe first consider the conditions in which machine learning may be especially useful with respect to two dimensions of a health outcome: (1) the characteristics of its diagnostic criteria and (2) the format in which its diagnostic data are usually stored within EHR systems. In the first dimension, we propose that for health outcomes with diagnostic criteria involving many clinical factors, vague definitions, or subjective interpretations, machine learning may be useful for modeling the complex diagnostic decision-making process from a vector of clinical inputs to identify individuals with the health outcome. In the second dimension, we propose that for health outcomes where diagnostic information is largely stored in unstructured formats such as free text or images, machine learning may be useful for extracting and structuring this information as part of a natural language processing system or an image recognition task. We then consider these two dimensions jointly to define four common scenarios of health outcomes. For each scenario, we discuss the potential uses for machine learning-first assuming accurate and complete EHR data and then relaxing these assumptions to accommodate the limitations of real-world EHR systems. We illustrate these four scenarios using concrete examples and describe how recent studies have used machine learning to identify these health outcomes from EHR data.SummaryMachine learning has great potential to improve the accuracy and efficiency of health outcome identification from EHR systems, especially under certain conditions. To promote the use of machine learning in EHR-based phenotyping tasks, future work should prioritize efforts to increase the transportability of machine learning algorithms for use in multi-site settings.