Recognizing clinical entities in hospital discharge summaries using Structural Support Vector Machines with word representation features.

Recognizing clinical entities in hospital discharge summaries using Structural Support Vector Machines with word representation features.
复制标题

DOI:
10.1186/1472-6947-13-s1-s1
复制
发表时间:
2013
影响因子:
3.5
通讯作者:
Xu H
Xu H
中科院分区:
医学3区
文献类型:
--
作者:
Tang B;Cao H;Wu Y;Jiang M;Xu H

文献摘要

被引文献

相似文献

命名实体识别(NER)是临床自然语言处理(NLP)研究的重要课题。基于机器学习(ML)的NER方法在识别临床文本中的实体方面表现出了良好的性能。算法和特征是影响基于ML的NER系统性能的两个重要因素。序贯标注算法条件随机场(CRF)和基于大间隔理论的支持向量机(SVMs)是两种典型的机器学习算法,已广泛应用于临床NER任务中。在特征方面,临床NER系统中经常使用语境词的句法和语义信息。然而,结构支持向量机(SSVMs)是一种结合了CRF和SVMs优点的算法,以及单词表示特征,这些特征通过非监督算法在大型未标记语料库上包含单词级别的后退信息,尚未被广泛研究用于临床文本处理。因此,本研究的主要目的是评估SSVMS和词语表征特征在临床NER任务中的使用情况。在这项研究中,我们开发了基于SSVMS的NER系统来识别医院出院摘要中的临床实体,使用了2010年i2b2 NLP挑战中概念提取任务的数据集。在特征集相同的情况下,我们比较了基于CRFS和SSVMS的NER分类器的性能。此外,我们还提取了两种不同类型的词表征特征(基于聚类的表征特征和分布表征特征),并将其集成到基于SSVMS的临床神经网络系统中。然后,我们报告了基于SSVM的具有不同类型的词表征特征的NER系统的性能。在挑战中使用相同的训练集(N=27,837)和测试集(N=45,009),我们的评估表明,当使用相同的特征时,基于SSVMS的NER系统在临床实体识别方面取得了比基于CRFS的系统更好的性能。两种类型的词表示特征(基于聚类的表示和分布表示)都提高了基于ML的NER系统的性能。通过将两种不同类型的词表征特征与SSVM相结合,我们的系统获得了85.82%的最高F-度量,比挑战中报道的最好的系统高出0.6%。实验结果表明,SSVMS是一种很有潜力的临床NLP研究算法,这两种无监督词语表征特征对临床NER任务都是有益的。
Named entity recognition (NER) is an important task in clinical natural language processing (NLP) research. Machine learning (ML) based NER methods have shown good performance in recognizing entities in clinical text. Algorithms and features are two important factors that largely affect the performance of ML-based NER systems. Conditional Random Fields (CRFs), a sequential labelling algorithm, and Support Vector Machines (SVMs), which is based on large margin theory, are two typical machine learning algorithms that have been widely applied to clinical NER tasks. For features, syntactic and semantic information of context words has often been used in clinical NER systems. However, Structural Support Vector Machines (SSVMs), an algorithm that combines the advantages of both CRFs and SVMs, and word representation features, which contain word-level back-off information over large unlabelled corpus by unsupervised algorithms, have not been extensively investigated for clinical text processing. Therefore, the primary goal of this study is to evaluate the use of SSVMs and word representation features in clinical NER tasks. In this study, we developed SSVMs-based NER systems to recognize clinical entities in hospital discharge summaries, using the data set from the concept extration task in the 2010 i2b2 NLP challenge. We compared the performance of CRFs and SSVMs-based NER classifiers with the same feature sets. Furthermore, we extracted two different types of word representation features (clustering-based representation features and distributional representation features) and integrated them with the SSVMs-based clinical NER system. We then reported the performance of SSVM-based NER systems with different types of word representation features. Using the same training (N = 27,837) and test (N = 45,009) sets in the challenge, our evaluation showed that the SSVMs-based NER systems achieved better performance than the CRFs-based systems for clinical entity recognition, when same features were used. Both types of word representation features (clustering-based and distributional representations) improved the performance of ML-based NER systems. By combining two different types of word representation features together with SSVMs, our system achieved a highest F-measure of 85.82%, which outperformed the best system reported in the challenge by 0.6%. Our results show that SSVMs is a great potential algorithm for clinical NLP research, and both types of unsupervised word representation features are beneficial to clinical NER tasks.