Automatic transliteration for Japanese-to-English text retrieval

Automatic transliteration for Japanese-to-English text retrieval
复制标题

用于日文到英文文本检索的自动音译

DOI:
10.1145/860435.860499
复制
发表时间:
2003
期刊:
Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval
影响因子:
--
通讯作者:
David A. Evans
David A. Evans
中科院分区:
--
文献类型:
--
作者:
Yan Qu;G. Grefenstette;David A. Evans

文献摘要

参考文献

被引文献

相似文献

对于基于双语翻译词典的跨语言信息检索(CLIR),良好的性能取决于词典中的词汇覆盖率。对于几乎没有同源语的语言来说尤其如此,例如日语和英语之间。在本文中,我们描述了一种自动创建和验证英语单词的候选日语音译术语的方法。使用语音英语词典和一组概率映射规则来自动生成音译候选。然后使用单语日语语料库自动验证音译术语。我们通过 CLEF 双语测试集上的日语到英语检索实验来评估提取的英语-日语音译对的使用情况。使用我们自动导出的双语翻译词典扩展,可以提高伪相关反馈之前和之后的平均精度,增益范围为 2.5% 到 64.8%。
For cross language information retrieval (CLIR) based on bilingual translation dictionaries, good performance depends upon lexical coverage in the dictionary. This is especially true for languages possessing few inter-language cognates, such as between Japanese and English. In this paper, we describe a method for automatically creating and validating candidate Japanese transliterated terms of English words. A phonetic English dictionary and a set of probabilistic mapping rules are used for automatically generating transliteration candidates. A monolingual Japanese corpus is then used for automatically validating the transliterated terms. We evaluate the usage of the extracted English-Japanese transliteration pairs with Japanese to English retrieval experiments over the CLEF bilingual test collections. The use of our automatically derived extension to a bilingual translation dictionary improves average precision, both before and after pseudo-relevance feedback, with gains ranging from 2.5% to 64.8%.
DOI: 10.1007/978-981-15-6168-9
发表时间: 2020-11
影响因子: 9.3
作者:
Robert E. Mercer
通讯作者: Robert E. Mercer