Joint segmentation and named entity recognition using dual decomposition in Chinese discharge summaries

Joint segmentation and named entity recognition using dual decomposition in Chinese discharge summaries
复制标题

中文出院小结中使用对偶分解的联合分割和命名实体识别

DOI:
10.1136/amiajnl-2013-001806
复制
发表时间:
2014-02-01
影响因子:
6.4
通讯作者:
Chang, Eric I.
Chang, Eric I.
中科院分区:
管理学2区
文献类型:
--
作者:
Xu, Yan;Wang, Yining;Chang, Eric I.

文献摘要

被引文献

相似文献

目的本文主要研究三个方面的内容:(1)建立一套中文出院小结标准语料库;(2)在标准语料库中进行分词和命名实体识别;(3)建立一个分词和命名实体识别的联合模型。在自然语言处理领域,虽然大多数方法使用单个模型来预测输出,但许多工作已经证明,可以通过利用组合技术来提高许多任务的性能。因此,在本文中,我们提出了一个联合模型,使用对偶分解来执行这两个任务,以利用这两个任务之间的相关性。结果建立了336个71355字的中文出院小结金标准语料库,并与独立模型、增量模型和基于组合标签的联合模型进行了比较。使用双重分解的框架实现了0.2%的改进分割和1%的改进识别,相比,每两个tasks.Conclusions联合模型是高效和有效的分割和识别相比,两个单独的任务。该模型取得了令人鼓舞的结果,证明了这两个任务的可行性。
Objective In this paper, we focus on three aspects: (1) to annotate a set of standard corpus in Chinese discharge summaries; (2) to perform word segmentation and named entity recognition in the above corpus; (3) to build a joint model that performs word segmentation and named entity recognition.Design Two independent systems of word segmentation and named entity recognition were built based on conditional random field models. In the field of natural language processing, while most approaches use a single model to predict outputs, many works have proved that performance of many tasks can be improved by exploiting combined techniques. Therefore, in this paper, we proposed a joint model using dual decomposition to perform both the two tasks in order to exploit correlations between the two tasks. Three sets of features were designed to demonstrate the advantage of the joint model we proposed, compared with independent models, incremental models and a joint model trained on combined labels.Measurements Micro-averaged precision (P), recall (R), and F-measure (F) were used to evaluate results.Results The gold standard corpus is created using 336 Chinese discharge summaries of 71 355 words. The framework using dual decomposition achieved 0.2% improvement for segmentation and 1% improvement for recognition, compared with each of the two tasks alone.Conclusions The joint model is efficient and effective in both segmentation and recognition compared with the two individual tasks. The model achieved encouraging results, demonstrating the feasibility of the two tasks.