The role of fine-grained annotations in supervised recognition of risk factors for heart disease from EHRs.

The role of fine-grained annotations in supervised recognition of risk factors for heart disease from EHRs.
复制标题

DOI:
10.1016/j.jbi.2015.06.010
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Demner-Fushman D
Demner-Fushman D
中科院分区:
医学3区
文献类型:
--
作者:
Roberts K;Shooshan SE;Rodriguez L;Abhyankar S;Kilicoglu H;Demner-Fushman D

文献摘要

被引文献

相似文献

本文描述了一种监督机器学习方法,用于识别临床文本中的心脏病风险因素,并评估注释粒度和质量对系统识别这些风险因素的能力的影响。我们利用一系列支持向量机模型与手动构建的词典相结合,对特定于每个风险因素的触发因素进行分类。用于分类的特征非常简单,仅利用词汇信息,而忽略语法和语义等更高级别的语言信息。相反,我们通过在标准语料库上注释附加信息来合并高质量数据来训练模型。尽管该系统相对简单,但它在 2014 年 i2b2/UTHealth 共享任务的 20 名参与者中获得了最高分(微观和宏观 F1,以及微观和宏观回忆)。该系统的微观(宏观)精度为 0.8951 (0.8965),召回率为 0.9625 (0.9611),F1 测量值为 0.9276 (0.9277)。此外,我们还进行了一系列实验来评估我们创建的注释数据的价值。这些实验展示了手动标记的负注释如何提高信息提取性能,证明了高质量、细粒度的自然语言注释的重要性。
This paper describes a supervised machine learning approach for identifying heart disease risk factors in clinical text, and assessing the impact of annotation granularity and quality on the system's ability to recognize these risk factors. We utilize a series of support vector machine models in conjunction with manually built lexicons to classify triggers specific to each risk factor. The features used for classification were quite simple, utilizing only lexical information and ignoring higher-level linguistic information such as syntax and semantics. Instead, we incorporated high-quality data to train the models by annotating additional information on top of a standard corpus. Despite the relative simplicity of the system, it achieves the highest scores (micro- and macro-F1, and micro- and macro-recall) out of the 20 participants in the 2014 i2b2/UTHealth Shared Task. This system obtains a micro- (macro-) precision of 0.8951 (0.8965), recall of 0.9625 (0.9611), and F1-measure of 0.9276 (0.9277). Additionally, we perform a series of experiments to assess the value of the annotated data we created. These experiments show how manually-labeled negative annotations can improve information extraction performance, demonstrating the importance of high-quality, fine-grained natural language annotations.