A Hybrid Model for Extracting Transliteration Equivalents from Parallel Corpora

A Hybrid Model for Extracting Transliteration Equivalents from Parallel Corpora
复制标题

从并行语料库中提取音译等价物的混合模型

DOI:
10.1007/11846406_15
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
H. Isahara
H. Isahara
中科院分区:
--
文献类型:
--
作者:
Jong;Key;H. Isahara

文献摘要

被引文献

相似文献

已经提出了几种音译对获取模型来克服音译引起的词汇外问题。然而,迄今为止,关于可以同时容纳多个模型的框架的文献还很少。此外,很少有人关心使用最新的语料库(例如网络文档)来验证获得的音译对。为了解决这些问题,我们提出了一种音译对获取的混合模型。在本文中,我们专注于一个结合多种音译对获取模型的框架。实验表明,我们的混合模型比单独的每个音译对获取模型更有效。
Several models for transliteration pair acquisition have been proposed to overcome the out-of-vocabulary problem caused by transliterations. To date, however, there has been little literature regarding a framework that can accommodate several models at the same time. Moreover, there is little concern for validating acquired transliteration pairs using up-to-date corpora, such as web documents. To address these problems, we propose a hybrid model for transliteration pair acquisition. In this paper, we concentrate on a framework for combining several models for transliteration pair acquisition. Experiments showed that our hybrid model was more effective than each individual transliteration pair acquisition model alone.