Exploring feature sets for two-phase biomedical named entity recognition using semi-CRFs

Exploring feature sets for two-phase biomedical named entity recognition using semi-CRFs
复制标题

DOI:
10.1007/s10115-013-0637-7
复制
发表时间:
2014-08
影响因子:
2.7
通讯作者:
Li Yang-;Yanhong Zhou
Li Yang-;Yanhong Zhou
中科院分区:
计算机科学4区
文献类型:
--
作者:
Li Yang-;Yanhong Zhou

文献摘要

被引文献

相似文献

本文提出了一种基于半马尔可夫条件随机场模型(semi-Markov conditional random fields model,semi-CRFs)的两阶段方法,并探索了新的特征集,用于将文本中的实体分为5种类型:蛋白质、DNA、RNA、细胞系和细胞类型。Semi-CRFs将标签放在一个片段上,而不是一个单词,这比其他机器学习方法,如条件随机场模型(CRFs)更自然。我们的方法将生物医学命名实体识别任务分为两个子任务:术语边界检测和语义标记。在第一阶段,术语边界检测子任务检测实体的边界,并将实体分类为一种类型C。在第二阶段,语义标记子任务将在第一阶段检测到的实体标记为正确的实体类型。我们在这两个阶段探索新的功能集,以提高性能。为了进行比较,在每个阶段对CRF和半CRF模型进行了实验。我们在JNLPBA 2004数据集上进行的实验在没有深度领域知识和后处理算法的情况下基于半CRF实现了74.64%的F分数,这优于大多数最先进的系统。
This paper represents a two-phase approach based on semi-Markov conditional random fields model (semi-CRFs) and explores novel feature sets for identifying the entities in text into 5 types: protein, DNA, RNA, cell_line and cell_type. Semi-CRFs put the label to a segment not a single word which is more natural than the other machine learning methods such as conditional random fields model (CRFs). Our approach divides the biomedical named entity recognition task into two sub-tasks: term boundary detection and semantic labeling. At the first phase, term boundary detection sub-task detects the boundary of the entities and classifies the entities into one type C. At the second phase, semantic labeling sub-task labels the entities detected at the first phase the correct entity type. We explore novel feature sets at both phases to improve the performance. To make a comparison, experiments conducted both on CRFs and on semi-CRFs models at each phase. Our experiments carried out on JNLPBA 2004 datasets achieve an F-score of 74.64 % based on semi-CRFs without deep domain knowledge and post-processing algorithms, which outperforms most of the state-of-the-art systems.