Deriving comorbidities from medical records using natural language processing

Deriving comorbidities from medical records using natural language processing
复制标题

DOI:
10.1136/amiajnl-2013-001889
复制
发表时间:
2013-12-01
影响因子:
6.4
通讯作者:
Friedman, Carol
Friedman, Carol
中科院分区:
管理学2区
文献类型:
--
作者:
Salmasian, Hojjat;Freedberg, Daniel E.;Friedman, Carol

文献摘要

被引文献

相似文献

由于共病的混杂效应,提取共病信息对于表型研究至关重要。我们开发了一种自动方法,可以准确地从电子病历中确定并存疾病。使用Charlson共病指数(CCI)的修改版本,两名医生通过手动审查100份入院记录创建了合并症的参考标准。我们使用MedLEE自然语言处理系统处理笔记,并编写查询以从其结构化输出中自动提取合并症。评价者之间对参考集的一致性非常高(97.7%)。我们的方法得到的F1得分为0.761,总的CCI得分与参考标准无差异(p=0.329,幂80.4%)。相比之下,由于敏感度较低(66.1%),从索赔数据获得的合并症F1得分为0.741。由于CCI先前已被证实是死亡率和再入院的预测因子,我们的方法可以自动预测这些结果。
Extracting comorbidity information is crucial for phenotypic studies because of the confounding effect of comorbidities. We developed an automated method that accurately determines comorbidities from electronic medical records. Using a modified version of the Charlson comorbidity index (CCI), two physicians created a reference standard of comorbidities by manual review of 100 admission notes. We processed the notes using the MedLEE natural language processing system, and wrote queries to extract comorbidities automatically from its structured output. Interrater agreement for the reference set was very high (97.7%). Our method yielded an F1 score of 0.761 and the summed CCI score was not different from the reference standard (p=0.329, power 80.4%). In comparison, obtaining comorbidities from claims data yielded an F1 score of 0.741, due to lower sensitivity (66.1%). Because CCI has previously been validated as a predictor of mortality and readmission, our method could allow automated prediction of these outcomes.