Building a Multilingual Lexical Resource for Named Entity Disambiguation, Translation and Transliteration

Building a Multilingual Lexical Resource for Named Entity Disambiguation, Translation and Transliteration
复制标题

为命名实体消歧、翻译和音译构建多语言词汇资源

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Matthias Hartung
Matthias Hartung
中科院分区:
--
文献类型:
--
作者:
Wolodja Wentland;Johannes Knopp;Carina Silberer;Matthias Hartung

文献摘要

被引文献

相似文献

在本文中,我们提出了 HeiNER,多语言海德堡命名实体资源。 HeiNER 包含 1,547,586 个已消除歧义的英语命名实体以及 15 种语言的翻译和音译。我们的工作建立在(Bunescu 和 Pasca,2006)中描述的方法的基础上,但将其扩展到多语言维度。通过利用在线百科全书维基百科中包含的跨语言信息,将命名实体翻译成各种目标语言。此外,HeiNER 为所有目标语言的每个 NE 提供语言上下文,这使其成为多语言命名实体识别、消歧和分类的宝贵资源。我们的评估结果与人类注释者的评估相比,我们从英语维基百科中提取的 NE 的精度高达 0.95。因此,这些源语言 NE 是我们多语言 NE 翻译方法的非常可靠的种子。
In this paper, we present HeiNER, the multilingual Heidelberg Named Entity Resource. HeiNER contains 1,547,586 disambiguated English Named Entities together with translations and transliterations to 15 languages. Our work builds on the approach described in (Bunescu and Pasca, 2006), yet extends it to a multilingual dimension. Translating Named Entities into the various target languages is carried out by exploiting crosslingual information contained in the online encyclopedia Wikipedia. In addition, HeiNER provides linguistic contexts for every NE in all target languages which makes it a valuable resource for multilingual Named Entity Recognition, Disambiguation and Classification. The results of our evaluation against the assessments of human annotators yield a high precision of 0.95 for the NEs we extract from the English Wikipedia. These source language NEs are thus very reliable seeds for our multilingual NE translation method.