Extracting principal diagnosis, co-morbidity and smoking status for asthma research: evaluation of a natural language processing system.

Extracting principal diagnosis, co-morbidity and smoking status for asthma research: evaluation of a natural language processing system.
复制标题

为哮喘研究提取主要诊断、并发症和吸烟状况:对自然语言处理系统的评估。

DOI:
10.1186/1472-6947-6-30
复制
发表时间:
2006-07-26
影响因子:
3.5
通讯作者:
Lazarus, Ross
Lazarus, Ross
中科院分区:
医学3区
文献类型:
--
作者:
Zeng, Qing T;Goryachev, Sergey;Lazarus, Ross

文献摘要

被引文献

相似文献

背景:电子病历中的文字描述是一个丰富的信息源。我们开发了一个健康信息文本提取(HITEX)工具,并使用它来提取气道疾病研究的关键发现。方法:将HITEX从一组150份出院摘要中提取的主要诊断、共病和吸烟状态与专家生成的金标准进行比较。当排除金标准标记为“数据不足”的病例时,HITEX的主要诊断准确率为82%,共病准确率为87%,吸烟状态提取准确率为90%。鉴于出院摘要和提取任务的复杂性,我们认为结果是有希望的。
BACKGROUND: The text descriptions in electronic medical records are a rich source of information. We have developed a Health Information Text Extraction (HITEx) tool and used it to extract key findings for a research study on airways disease.METHODS: The principal diagnosis, co-morbidity and smoking status extracted by HITEx from a set of 150 discharge summaries were compared to an expert-generated gold standard.RESULTS: The accuracy of HITEx was 82% for principal diagnosis, 87% for co-morbidity, and 90% for smoking status extraction, when cases labeled "Insufficient Data" by the gold standard were excluded.CONCLUSION: We consider the results promising, given the complexity of the discharge summaries and the extraction tasks.