Automated concept-level information extraction to reduce the need for custom software and rules development

Automated concept-level information extraction to reduce the need for custom software and rules development
复制标题

DOI:
10.1136/amiajnl-2011-000183
复制
发表时间:
2011-09-01
影响因子:
6.4
通讯作者:
Fiore, Louis D.
Fiore, Louis D.
中科院分区:
管理学2区
文献类型:
--
作者:
D'Avolio, Leonard W.;Nguyen, Thien M.;Fiore, Louis D.

文献摘要

被引文献

相似文献

尽管至少有40年的经验表现,但目前很少有临床自然语言处理(NLP)或信息提取系统对医学科学或护理做出贡献。作者通过使用图形用户界面驱动的、高度可概括的概念级检索方法来减少对定制软件和规则开发的需求,从而解决了这一差距。材料和方法“通过示例学习”方法将来自开源NLP管道的功能与开源机器学习分类器相结合,以自动和迭代地评估最佳性能配置。第四届i2 b2/VA共享任务挑战赛的概念提取任务提供的数据集和指标,用于评估performance.Results的F-测量得分为每个任务的医疗问题(0.83),治疗(0.82),和测试(0.83)。在所有的实验中,回忆都滞后于精确。所有任务的精度都接近或高于0.90。讨论在没有任务定制和最终用户配置和启动每个实验的时间不到5分钟的情况下,平均F-测量值为0.83,比比赛中22名参赛者的平均F-测量值落后一个点。强大的精度分数表明应用该方法进行更具体的临床信息提取任务的潜力。有没有一个最好的配置,支持一个迭代的方法模型creation.Conclusion性能可以接受的水平,可以实现使用完全自动化和概括的方法,概念级的信息提取。所描述的实现和相关文档可供下载。
Objective Despite at least 40 years of promising empirical performance, very few clinical natural language processing (NLP) or information extraction systems currently contribute to medical science or care. The authors address this gap by reducing the need for custom software and rules development with a graphical user interface-driven, highly generalizable approach to concept-level retrieval.Materials and methods A 'learn by example' approach combines features derived from open-source NLP pipelines with open-source machine learning classifiers to automatically and iteratively evaluate top-performing configurations. The Fourth i2b2/VA Shared Task Challenge's concept extraction task provided the data sets and metrics used to evaluate performance.Results Top F-measure scores for each of the tasks were medical problems (0.83), treatments (0.82), and tests (0.83). Recall lagged precision in all experiments. Precision was near or above 0.90 in all tasks.Discussion With no customization for the tasks and less than 5 min of end-user time to configure and launch each experiment, the average F-measure was 0.83, one point behind the mean F-measure of the 22 entrants in the competition. Strong precision scores indicate the potential of applying the approach for more specific clinical information extraction tasks. There was not one best configuration, supporting an iterative approach to model creation.Conclusion Acceptable levels of performance can be achieved using fully automated and generalizable approaches to concept-level information extraction. The described implementation and related documentation is available for download.