Automatically extracting information needs from complex clinical questions.

Automatically extracting information needs from complex clinical questions.
复制标题

DOI:
10.1016/j.jbi.2010.07.007
复制
发表时间:
2010-12
影响因子:
4.5
通讯作者:
Yu H
Yu H
中科院分区:
医学3区
文献类型:
--
作者:
Cao YG;Cimino JJ;Ely J;Yu H

文献摘要

参考文献

被引文献

相似文献

临床医生在看病人时提出复杂的临床问题,及时确定这些问题的答案有助于提高病人护理的质量。本文报道了两种自然语言处理模型,即自动主题分配和关键词识别,它们可以自动有效地从特殊临床问题中提取信息需求。我们的研究是在开发更大的临床问答系统AskHERMES(帮助临床医生提取和阐明多媒体信息以回答临床问题)的背景下进行的。我们开发了有监督的机器学习系统,为问题自动分配预定义的一般类别(例如,病因、程序和诊断)。我们还探索了有监督和无监督系统,以自动识别捕获问题主要内容的关键字。我们根据实践中收集的4,654个带注释的临床问题评估了我们的系统。我们在一般主题分类任务中获得了76.0%的F1分数,在关键词提取任务中获得了58.0%的F1分数。我们的系统已经被应用到更大的问答系统AskHERMES中。我们的错误分析表明,在我们的训练数据中不一致的注释已经损害了两个问题分析任务。我们的系统可以在http://www.askhermes.org上获得,它可以自动从短问题(单词令牌的数量<20)和长问题(单词令牌的数量bbb20)以及结构良好和格式不良的问题中提取信息需求。我们推测,如果提供一致注释的数据,则可以进一步提高一般主题分类和关键字提取的性能。
Clinicians pose complex clinical questions when seeing patients, and identifying the answers to those questions in a timely manner helps improve the quality of patient care. We report here on two natural language processing models, namely, automatic topic assignment and keyword identification, that together automatically and effectively extract information needs from ad hoc clinical questions. Our study is motivated in the context of developing the larger clinical question answering system AskHERMES (Help clinicians to Extract and aRrticulate Multimedia information for answering clinical quEstionS). We developed supervised machine-learning systems to automatically assign predefined general categories (e.g., etiology, procedure, and diagnosis) to a question. We also explored both supervised and unsupervised systems to automatically identify keywords that capture the main content of the question. We evaluated our systems on 4,654 annotated clinical questions that were collected in practice. We achieved an F1 score of 76.0% for the task of general topic classification and 58.0% for keyword extraction. Our systems have been implemented into the larger question answering system AskHERMES. Our error analyses suggested that inconsistent annotation in our training data have hurt both question analysis tasks. Our systems, available at http://www.askhermes.org, can automatically extract information needs from both short (the number of word tokens <20) and long questions (the number of word tokens >20), and from both well-structured and ill-formed questions. We speculate that the performance of general topic classification and keyword extraction can be further improved if consistently annotated data are made available.
DOI: 10.7326/0003-4819-103-4-596
发表时间: 1985-01-01
影响因子: 39.2
作者:
COVELL, DG;UMAN, GC;MANNING, PR
通讯作者: MANNING, PR
DOI: 10.1162/coli.2007.33.1.63
发表时间: 2007-03-01
影响因子: 9.3
作者:
Demner-Fushman, Dina;Lin, Jimmy
通讯作者: Lin, Jimmy
DOI: 10.1136/bmj.319.7206.358
发表时间: 1999-08-07
影响因子: --
作者:
Ely, JW;Osheroff, JA;Evans, ER
通讯作者: Evans, ER
DOI: 10.1136/bmj.321.7258.429
发表时间: 2000-08-12
影响因子: --
作者:
Ely, JW;Osheroff, JA;Stavri, PZ
通讯作者: Stavri, PZ
DOI: 10.1136/bmj.324.7339.710
发表时间: 2002-03-23
影响因子: --
作者:
Ely, JW;Osheroff, JA;Pifer, EA
通讯作者: Pifer, EA