Assessing the role of a medication-indication resource in the treatment relation extraction from clinical text

Assessing the role of a medication-indication resource in the treatment relation extraction from clinical text
复制标题

DOI:
10.1136/amiajnl-2014-002954
复制
发表时间:
2015-04-01
影响因子:
6.4
通讯作者:
Denny, Joshua C.
Denny, Joshua C.
中科院分区:
管理学2区
文献类型:
--
作者:
Bejan, Cosmin Adrian;Wei, Wei-Qi;Denny, Joshua C.

文献摘要

被引文献

相似文献

目的评价药物适应证(MEDI)资源和SemRep在临床文本中识别治疗关系的作用。材料与方法首先利用SemRep对临床文献进行处理,提取统一医学语言系统(UMLS)概念及其之间的治疗关系。然后,我们将MEDI合并到一个简单的算法中,该算法识别两个概念之间的治疗关系,如果它们与该资源中的药物-适应症对匹配。为了更好地覆盖范围,我们使用RxNorm和UMLS Metathesaurus的本体关系扩展了MEDI。我们还开发了两种集成方法,将SemRep的预测和MEDI算法相结合。我们在两个数据集上对我们选择的方法进行了评估,一个是Vanderbilt语料库,包含6864个出院摘要,另一个是2010年生物学与床边整合信息学(I2b2)/退伍军人事务(VA)挑战数据集。对25%的关系进行双重注释,符合率较高(Cohen‘s kappa=0.86)。评估包括将手动注释的关系与SemRep、MEDI算法和两种集成方法确定的关系进行比较。在第一个数据集上,MEDI算法和两个资源的联合获得的最佳F1测量结果(分别为78.7和80)显著高于SemRep的结果(72.3)。在第二个数据集上,与i2b2挑战中最好的系统相比,MEDI算法获得了更高的准确率和明显更低的召回值。这两个系统在i2b2关系子集上获得了与MEDI中的两个参数都具有可比性的F1测量值。结论SemRep和MEDI都可以用于从临床文本中提取治疗关系。使用MEDI的基于知识的提取优于单独使用SemRep,但通过集成这两个系统获得了更好的性能。将MEDI等基于知识的资源整合到SemRep和i2b2关系抽取器等信息抽取系统中,可能会改进从临床文本中提取治疗关系。
Objective To evaluate the contribution of the MEDication Indication (MEDI) resource and SemRep for identifying treatment relations in clinical text.Materials and methods We first processed clinical documents with SemRep to extract the Unified Medical Language System (UMLS) concepts and the treatment relations between them. Then, we incorporated MEDI into a simple algorithm that identifies treatment relations between two concepts if they match a medication-indication pair in this resource. For a better coverage, we expanded MEDI using ontology relationships from RxNorm and UMLS Metathesaurus. We also developed two ensemble methods, which combined the predictions of SemRep and the MEDI algorithm. We evaluated our selected methods on two datasets, a Vanderbilt corpus of 6864 discharge summaries and the 2010 Informatics for Integrating Biology and the Bedside (i2b2)/Veteran's Affairs (VA) challenge dataset.Results The Vanderbilt dataset included 958 manually annotated treatment relations. A double annotation was performed on 25% of relations with high agreement (Cohen's kappa = 0.86). The evaluation consisted of comparing the manual annotated relations with the relations identified by SemRep, the MEDI algorithm, and the two ensemble methods. On the first dataset, the best F1-measure results achieved by the MEDI algorithm and the union of the two resources (78.7 and 80, respectively) were significantly higher than the SemRep results (72.3). On the second dataset, the MEDI algorithm achieved better precision and significantly lower recall values than the best system in the i2b2 challenge. The two systems obtained comparable F1-measure values on the subset of i2b2 relations with both arguments in MEDI.Conclusions Both SemRep and MEDI can be used to extract treatment relations from clinical text. Knowledge-based extraction with MEDI outperformed use of SemRep alone, but superior performance was achieved by integrating both systems. The integration of knowledge-based resources such as MEDI into information extraction systems such as SemRep and the i2b2 relation extractors may improve treatment relation extraction from clinical text.