Robust Extraction of Named Entity Including Unfamiliar Word

Robust Extraction of Named Entity Including Unfamiliar Word
复制标题

DOI:
10.3115/1557690.1557723
复制
发表时间:
2008-06
期刊:
--
影响因子:
--
通讯作者:
Masatoshi Tsuchiya;Shinya Hida;S. Nakagawa
Masatoshi Tsuchiya;Shinya Hida;S. Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Masatoshi Tsuchiya;Shinya Hida;S. Nakagawa

文献摘要

相似文献

本文提出了一种新的方法来提取命名实体,包括不熟悉的词,不出现或出现在一个训练语料库中的一个大型的未标注的语料库。所提出的方法包括两个步骤。第一步是根据从大型未注释语料库中计算出的上下文向量,将最相似和最熟悉的单词分配给每个不熟悉的单词。然后,采用传统的机器学习方法作为第二步。通过对IREX语料库和NHK语料库的日语命名实体抽取实验,验证了该方法的有效性。
This paper proposes a novel method to extract named entities including unfamiliar words which do not occur or occur few times in a training corpus using a large unannotated corpus. The proposed method consists of two steps. The first step is to assign the most similar and familiar word to each unfamiliar word based on their context vectors calculated from a large unannotated corpus. After that, traditional machine learning approaches are employed as the second step. The experiments of extracting Japanese named entities from IREX corpus and NHK corpus show the effectiveness of the proposed method.