A sequence labeling approach to link medications and their attributes in clinical notes and clinical trial announcements for information extraction.

A sequence labeling approach to link medications and their attributes in clinical notes and clinical trial announcements for information extraction.
复制标题

DOI:
10.1136/amiajnl-2012-001487
复制
发表时间:
2013-09
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Solti I
Solti I
中科院分区:
其他
文献类型:
--
作者:
Li Q;Zhai H;Deleger L;Lingren T;Kaiser M;Stoutenborough L;Solti I

文献摘要

参考文献

被引文献

相似文献

这项工作的目标是评估机器学习方法、二元分类和序列标记,用于两个临床语料库中的药物属性连锁检测。我们对药物命名实体及其属性的 3000 个临床试验公告 (CTA) 和 1655 个临床注释 (CN) 进行了双重注释。提出了一种具有简约特征集的二元支持向量机(SVM)分类方法和基于条件随机场(CRF)的多层序列标记(MLSL)模型来识别实体及其相应属性之间的联系。我们根据人类制定的黄金标准评估了系统的性能。实验表明,这两种机器学习方法的性能在统计上显着优于基于规则的基线方法。以单个标记作为特征,二元 SVM 分类实现了 0.94 F 测量。在简约特征集上训练的 SVM 模型,CN 的 F 测量值达到 0.81,CTA 的 F 测量值达到 0.87。 CRF MLSL 方法在两个语料库上实现了 0.80 F 测量。我们将新颖的 MLSL 方法与二元分类和基于规则的方法进行了比较。 MLSL 方法在统计上的表现明显优于基于规则的方法。然而,对于 CTA 和 CN 语料库,基于 SVM 的二元分类方法在统计上显着优于 MLSL 方法。使用简约的特征集,基于 SVM 的二元分类和基于 CRF 的 MLSL 方法在检测 CTA 和 CN 中的药物名称和属性链接方面实现了高性能。
The goal of this work was to evaluate machine learning methods, binary classification and sequence labeling, for medication–attribute linkage detection in two clinical corpora. We double annotated 3000 clinical trial announcements (CTA) and 1655 clinical notes (CN) for medication named entities and their attributes. A binary support vector machine (SVM) classification method with parsimonious feature sets, and a conditional random fields (CRF)-based multi-layered sequence labeling (MLSL) model were proposed to identify the linkages between the entities and their corresponding attributes. We evaluated the system's performance against the human-generated gold standard. The experiments showed that the two machine learning approaches performed statistically significantly better than the baseline rule-based approach. The binary SVM classification achieved 0.94 F-measure with individual tokens as features. The SVM model trained on a parsimonious feature set achieved 0.81 F-measure for CN and 0.87 for CTA. The CRF MLSL method achieved 0.80 F-measure on both corpora. We compared the novel MLSL method with a binary classification and a rule-based method. The MLSL method performed statistically significantly better than the rule-based method. However, the SVM-based binary classification method was statistically significantly better than the MLSL method for both the CTA and CN corpora. Using parsimonious feature sets both the SVM-based binary classification and CRF-based MLSL methods achieved high performance in detecting medication name and attribute linkages in CTA and CN.
DOI: 10.1136/amiajnl-2011-000302
发表时间: 2011-09-01
影响因子: 6.4
作者:
Patrick, Jon D.;Nguyen, Dung H. M.;Li, Min
通讯作者: Li, Min
DOI: 10.1136/amiajnl-2011-000183
发表时间: 2011-09-01
影响因子: 6.4
作者:
D'Avolio, Leonard W.;Nguyen, Thien M.;Fiore, Louis D.
通讯作者: Fiore, Louis D.
DOI: 10.1136/jamia.2010.003855
发表时间: 2010-09-01
影响因子: 6.4
作者:
Doan, Son;Bastarache, Lisa;Xu, Hua
通讯作者: Xu, Hua
DOI: 10.1136/jamia.2010.004036
发表时间: 2010-09-01
影响因子: 6.4
作者:
Hamon, Thierry;Grabar, Natalia
通讯作者: Grabar, Natalia
DOI: 10.1136/jamia.2010.003962
发表时间: 2010-09-01
影响因子: 6.4
作者:
Deleger, Louise;Grouin, Cyril;Zweigenbaum, Pierre
通讯作者: Zweigenbaum, Pierre