Agile text mining for the 2014 i2b2/UTHealth Cardiac risk factors challenge

Agile text mining for the 2014 i2b2/UTHealth Cardiac risk factors challenge
复制标题

DOI:
10.1016/j.jbi.2015.06.030
复制
发表时间:
2015-12-01
影响因子:
4.5
通讯作者:
Jonnalagadda, Siddhartha R.
Jonnalagadda, Siddhartha R.
中科院分区:
医学3区
文献类型:
--
作者:
Cormack, James;Nath, Chinmoy;Jonnalagadda, Siddhartha R.

文献摘要

被引文献

相似文献

本文描述了使用敏捷文本挖掘平台(Linguamatics的交互式信息提取平台,12 E)来提取i2 b2/UTHealth 2014挑战中定义的患者记录中的文档级心脏风险因素。该方法使用了一个数据驱动的基于规则的方法,增加了一个简单的监督分类。我们证明了敏捷文本挖掘允许快速优化提取策略,而后处理可以利用注释指南,语料库统计数据和从黄金标准数据推断的逻辑。我们还展示了训练集中的数据不平衡如何影响性能。在测试数据上对这种方法的评估给出了91.7%的F分数,比性能最好的系统落后1%。(C)2015爱思唯尔公司All rights reserved.
This paper describes the use of an agile text mining platform (Linguamatics' Interactive Information Extraction Platform, 12E) to extract document-level cardiac risk factors in patient records as defined in the i2b2/UTHealth 2014 challenge. The approach uses a data-driven rule-based methodology with the addition of a simple supervised classifier. We demonstrate that agile text mining allows for rapid optimization of extraction strategies, while post-processing can leverage annotation guidelines, corpus statistics and logic inferred from the gold standard data. We also show how data imbalance in a training set affects performance. Evaluation of this approach on the test data gave an F-Score of 91.7%, one percent behind the top performing system. (C) 2015 Elsevier Inc. All rights reserved.