Subcategorizing EHR diagnosis codes to improve clinical application of machine learning models.

Subcategorizing EHR diagnosis codes to improve clinical application of machine learning models.
复制标题

DOI:
10.1016/j.ijmedinf.2021.104588
复制
发表时间:
2021-12
影响因子:
4.9
通讯作者:
Koroukian SM
Koroukian SM
中科院分区:
医学2区
文献类型:
--
作者:
Reimer AP;Dai W;Smith B;Schiltz NK;Sun J;Koroukian SM

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)数据通常用于次要目的,如研究和临床决策支持。然而,EHR数据的重用提出了若干挑战,包括但不限于识别与患者的临床遭遇相关联的所有诊断。本研究的目的是评估的可行性,开发一个模式,以确定和细分所有结构化的诊断代码的病人遇到。为了开发子分类模式,我们使用了来自医院间运输数据存储库的EHR数据,其中包含完整的医院就诊水平数据。八个离散的数据源,包含结构化的诊断代码被确定。使用统一医学语言系统对诊断代码进行标准化,并将额外的EHR数据与标准化术语相结合,以创建和验证子类别。然后,我们采用随机森林来评估新的亚分类诊断的有用性,以预测医院间转移后死亡率,通过建立2个模型,一个使用标准的诊断代码,一个使用新的亚分类诊断代码。诊断的六个子类别被确定和验证。这些子类别包括:原发或入院诊断(10%),既往医疗、手术或社会史(9%),问题列表(20%),合并症(24%),出院诊断(6%)和未映射诊断(31%)。亚分类模型的表现优于标准模型,训练AUROC为0.97对0.95,测试模型AUROC为0.81对0.46。我们的工作表明,将结构化诊断代码与额外的EHR数据和二级数据源合并,可以提供额外的信息来了解诊断在整个临床过程中的作用,并提高预测模型的性能。进一步的工作是必要的,以评估是否亚分类产生的好处,在解释预后模型的结果和/或操作的结果在临床决策支持应用。
Electronic health record (EHR) data is commonly used for secondary purposes such as research and clinical decision support. However, reuse of EHR data presents several challenges including but not limited to identifying all diagnoses associated with a patient’s clinical encounter. The purpose of this study was to assess the feasibility of developing a schema to identify and subclassify all structured diagnosis codes for a patient encounter. To develop a subclassification schema we used EHR data from an interhospital transport data repository that contained complete hospital encounter level data. Eight discrete data sources containing structured diagnosis codes were identified. Diagnosis codes were normalized using the Unified Medical Language System and additional EHR data were combined with standardized terminologies to create and validate the subcategories. We then employed random forest to assess the usefulness of the new subcategorized diagnoses to predict post-interhospital transfer mortality by building 2 models, one using standard diagnosis codes, and one using the new subcategorized diagnosis codes. Six subcategories of diagnoses were identified and validated. The subcategories included: primary or admitting diagnoses (10%), past medical, surgical or social history (9%), problem list (20%), comorbidity (24%), discharge diagnoses (6%), and unmapped diagnoses (31%). The subcategorized model outperformed the standard model, achieving a training AUROC of 0.97 versus 0.95 and testing model AUROC of 0.81 versus 0.46. Our work demonstrates that merging structured diagnosis codes with additional EHR data and secondary data sources provides additional information to understand the role of diagnosis throughout a clinical encounter and improves predictive model performance. Further work is necessary to assess if subcategorizing produces benefits in interpreting the results of prognostic models and/or operationalizing the results in clinical decision support applications.
DOI: 10.1371/journal.pone.0116656
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Mate S;Köpcke F;Toddenroth D;Martin M;Prokosch HU;Bürkle T;Ganslandt T
通讯作者: Ganslandt T
DOI: 10.1371/journal.pone.0246669
发表时间: 2021
期刊: PloS one
影响因子: 3.7
作者:
Altieri Dunn SC;Bellon JE;Bilderback A;Borrebach JD;Hodges JC;Wisniewski MK;Harinstein ME;Minnier TE;Nelson JB;Hall DE
通讯作者: Hall DE
DOI: 10.1002/cpt.1226
发表时间: 2019-04
影响因子: 6.7
作者:
Eichler HG;Bloechl-Daum B;Broich K;Kyrle PA;Oderkirk J;Rasi G;Santos Ivo R;Schuurman A;Senderovitz T;Slawomirski L;Wenzl M;Paris V
通讯作者: Paris V
DOI: 10.1097/00005650-199801000-00004
发表时间: 1998-01-01
期刊: MEDICAL CARE
影响因子: 3
作者:
Elixhauser, A;Steiner, C;Coffey, RN
通讯作者: Coffey, RN
DOI: 10.4037/ajcc2013426
发表时间: 2013-07-01
影响因子: 2.7
作者:
Collins, Sarah A.;Cato, Kenrick;Vawdrey, David K.
通讯作者: Vawdrey, David K.