Classifying free-text triage chief complaints into syndromic categories with natural language processing

Classifying free-text triage chief complaints into syndromic categories with natural language processing
复制标题

DOI:
10.1016/j.artmed.2004.04.001
复制
发表时间:
2005-01-01
影响因子:
7.5
通讯作者:
Olszewski, RT
Olszewski, RT
中科院分区:
工程技术1区
文献类型:
--
作者:
Chapman, WW;Christensen, LM;Olszewski, RT

文献摘要

被引文献

相似文献

目标:开发和评估自然语言处理应用程序,用于将主诉分类为症状类别以进行症状监测。简介:医疗领域人工智能应用的输入数据大部分是自由文本的患者病历,包括口述医疗报告和分诊主诉。为了对自动化系统有用,自由文本必须转换为编码形式。方法:我们实施了宾夕法尼亚州的生物监视检测系统来监测 2002 年冬季奥运会。由于输入数据采用自由文本格式,因此我们使用自然语言处理文本分类器将自由文本分类主诉自动分类为生物监测系统使用的症状类别。该分类器接受了来自宾夕法尼亚州的 4700 份主要投诉的训练。我们使用来自犹他州的 800 个主诉测试集来评估分类器将自由文本主诉分类为综合症类别的能力。结果:分类器在 ROC 曲线下产生以下面积:体质 = 0.95;胃肠道=0.97;出血=0.99;神经学 = 0.96;皮疹=1.0;呼吸=0.99;其他 = 0.96。使用系统语义模型中存储的信息,我们从呼吸分类中提取下呼吸道疾病和发烧下呼吸道疾病,精度分别为 0.97 和 0.96。结论:结果表明,可训练的自然语言处理文本分类器可以准确地从自由文本主诉中提取数据以进行生物监测。 (C) 2004 Elsevier B.V. 保留所有权利。
Objective: Develop and evaluate a natural language processing application for classifying chief complaints into syndromic categories for syndromic surveillance. Introduction: Much of the input data for artificial intelligence applications in the medical field are free-text patient medical records, including dictated medical reports and triage chief complaints. To be useful for automated systems, the free-text must be translated into encoded form. Methods: We implemented a biosurveillance detection system from Pennsylvania to monitor the 2002 Winter Olympic Games. Because input data was in free-text format, we used a natural language processing text classifier to automatically classify free-text triage chief complaints into syndromic categories used by the biosurveillance system. The classifier was trained on 4700 chief complaints from Pennsylvania. We evaluated the ability of the classifier to classify free-text chief complaints into syndromic categories with a test set of 800 chief complaints from Utah. Results: The classifier produced the following areas under the ROC curve: Constitutional = 0.95; Gastrointestinal = 0.97; Hemorrhagic = 0.99; Neurological = 0.96; Rash = 1.0; Respiratory = 0.99; Other = 0.96. Using information stored in the system's semantic model, we extracted from the Respiratory classifications lower respiratory complaints and lower respiratory complaints with fever with a precision of 0.97 and 0.96, respectively. Conclusion: Results suggest that a trainable natural Language processing text classifier can accurately extract data from free-text chief complaints for biosurveillance. (C) 2004 Elsevier B.V. All rights reserved.