Identifying patient smoking status from medical discharge records

Identifying patient smoking status from medical discharge records
复制标题

DOI:
10.1197/jamia.m2408
复制
发表时间:
2008-01-01
影响因子:
6.4
通讯作者:
Kohane, Isaac
Kohane, Isaac
中科院分区:
管理学2区
文献类型:
--
作者:
Uzuner, Oezlem;Goldstein, Ira;Kohane, Isaac

文献摘要

被引文献

相似文献

作者组织了一场自然语言处理 (NLP) 挑战赛,旨在根据出院记录中的信息自动确定患者的吸烟状况。该挑战赛是作为 i2b2(将生物学整合到床边的信息学)项目的一部分发布的,旨在调查、促进和检查临床叙述的医学语言理解研究。本文描述了吸烟挑战,详细介绍了数据和注释过程,解释了评估指标,讨论了为挑战开发的系统的特征,对收到的系统运行结果进行了分析,得出了有关现有技术的结论,并确定了未来研究的方向。共有11支队伍参加了吸烟挑战赛。每个团队最多提交 3 个系统运行,总共提供 23 个提交。使用微平均和宏平均精度、召回率和 F 测量对提交的系统运行进行评估。接受吸烟挑战的系统代表了各种机器学习和基于规则的算法。尽管他们的吸烟状况识别方法存在差异,但其中许多系统都提供了良好的结果。有 12 次系统运行的微平均 F 测量值高于 0.84。对结果的分析强调了这样一个事实:出院摘要使用有限数量的文本特征(例如“吸烟”、“烟草”、“雪茄”、社会历史等)来表达吸烟状况。许多有效的吸烟状况识别器都受益于这些功能。
The authors organized a Natural Language Processing (NLP) challenge on automatically determining the smoking status of patients from information found in their discharge records. This challenge was issued as a part of the i2b2 (Informatics for Integrating Biology to the Bedside) project, to survey, facilitate, and examine studies in medical language understanding for clinical narratives. This article describes the smoking challenge, details the data and the annotation process, explains the evaluation metrics, discusses the characteristics of the systems developed for the challenge, presents an analysis of the results of received system runs, draws conclusions about the state of the art, and identifies directions for future research. A total of 11 teams participated in the smoking challenge. Each team submitted up to three system runs, providing a total of 23 submissions. The submitted system runs were evaluated with microaveraged and macroaveraged precision, recall, and F-measure. The systems submitted to the smoking challenge represented a variety of machine learning and rule-based algorithms. Despite the differences in their approaches to smoking status identification, many of these systems provided good results. There were 12 system runs with microaveraged F-measures above 0.84. Analysis of the results highlighted the fact that discharge summaries express smoking status using a limited number of textual features (e.g., "smok", "tobac", "cigar", Social History, etc.). Many of the effective smoking status identifiers benefit from these features.