An empirical evaluation of supervised learning approaches in assigning diagnosis codes to electronic medical records.

An empirical evaluation of supervised learning approaches in assigning diagnosis codes to electronic medical records.
复制标题

DOI:
10.1016/j.artmed.2015.04.007
复制
发表时间:
2015-10
影响因子:
7.5
通讯作者:
Lu Y
Lu Y
中科院分区:
工程技术1区
文献类型:
--
作者:
Kavuluru R;Rios A;Lu Y

文献摘要

被引文献

相似文献

诊断代码由训练有素的编码员通过查看与患者来访相关的所有医生编写的文档来分配给医疗机构中的医疗记录。这是一项必要而复杂的任务,涉及编码员遵守编码准则并对所有可分配代码进行编码。近年来,随着电子病历(EMRS)的普及,编码分配的计算方法被提出。然而,大多数努力都集中在单一且往往简短的临床叙述上,而现实情况下需要对代码分配进行全面的EMR水平分析。我们评估了有监督的学习方法,以自动分配国际疾病分类(第九版)-临床修改(ICD-9-CM)代码给EMR,通过实验使用大量现实的EMR数据集。总体目标是确定在考虑此类数据集时在此任务中提供卓越性能的方法。我们使用了来自肯塔基大学(UKY)医学中心的71,463名EMR的数据集,这些EMR对应于出院日期在两年内(2011-2012)的住院患者。我们整理了这个数据集的一个较小的子集,也使用了第三个放射学报告的黄金标准数据集。我们使用不同的问题转换方法对特征和数据选择组件进行实验,并使用适当的标签校准和排序方法使用新的特征,包括代码共现频率和潜在代码关联。在包含至少50个训练样本的所有代码上,我们获得了0.48的微F分数。在至少出现在两年数据集的1%的代码集上,我们获得了0.54的微F分数。对于较小的放射学报告数据集,分类器链接方法会产生最好的结果。对于UKY数据集的较小子集,特征选择、数据选择和标签校准提供了最佳性能。我们表明,不同规模(EMR的大小、不同码数)和不同特征的数据集需要不同的学习方法。对于属于特定医学子领域(例如,放射学、病理学)的较短的叙述,鉴于代码彼此高度相关,分类器链接是理想的。对于现实的住院完整EMR,特征和数据选择方法为较小的数据集提供了高性能。然而,对于大型电子病历数据集,我们观察到具有基于学习到排名的代码重新排序的二进制相关性方法提供了最佳性能。无论训练数据集的大小如何,对于一般的EMRS来说,选择最优标签数量的标签校准是必不可少的最后一步。
Diagnosis codes are assigned to medical records in healthcare facilities by trained coders by reviewing all physician authored documents associated with a patient's visit. This is a necessary and complex task involving coders adhering to coding guidelines and coding all assignable codes. With the popularity of electronic medical records (EMRs), computational approaches to code assignment have been proposed in the recent years. However, most efforts have focused on single and often short clinical narratives, while realistic scenarios warrant full EMR level analysis for code assignment. We evaluate supervised learning approaches to automatically assign international classification of diseases (ninth revision) - clinical modification (ICD-9-CM) codes to EMRs by experimenting with a large realistic EMR dataset. The overall goal is to identify methods that offer superior performance in this task when considering such datasets. We use a dataset of 71,463 EMRs corresponding to in-patient visits with discharge date falling in a two year period (2011–2012) from the University of Kentucky (UKY) Medical Center. We curate a smaller subset of this dataset and also use a third gold standard dataset of radiology reports. We conduct experiments using different problem transformation approaches with feature and data selection components and employing suitable label calibration and ranking methods with novel features involving code co-occurrence frequencies and latent code associations. Over all codes with at least 50 training examples we obtain a micro F-score of 0.48. On the set of codes that occur at least in 1% of the two year dataset, we achieve a micro F-score of 0.54. For the smaller radiology report dataset, the classifier chaining approach yields best results. For the smaller subset of the UKY dataset, feature selection, data selection, and label calibration offer best performance. We show that datasets at different scale (size of the EMRs, number of distinct codes) and with different characteristics warrant different learning approaches. For shorter narratives pertaining to a particular medical subdomain (e.g., radiology, pathology), classifier chaining is ideal given the codes are highly related with each other. For realistic in-patient full EMRs, feature and data selection methods offer high performance for smaller datasets. However, for large EMR datasets, we observe that the binary relevance approach with learning-to-rank based code reranking offers the best performance. Regardless of the training dataset size, for general EMRs, label calibration to select the optimal number of labels is an indispensable final step.