Hypothesis Selection in Machine Transliteration: A Web Mining Approach

Hypothesis Selection in Machine Transliteration: A Web Mining Approach
复制标题

机器音译中的假设选择:一种网络挖掘方法

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
H. Isahara
H. Isahara
中科院分区:
--
文献类型:
--
作者:
Jong;H. Isahara

文献摘要

被引文献

相似文献

提出了一种选择机器音译假设的新方法。我们为给定的英语单词生成一组中文、日文和韩文的音译假设。然后,我们使用一组音译假设作为查找相关Web页面和从Web页面中为音译假设挖掘上下文信息的指南。最后,我们将挖掘的信息用于机器学习算法,包括支持向量机和最大熵模型,以选择正确的音译假设。在我们的实验中,我们提出的基于Web挖掘的方法始终优于以前工作中使用的基于简单Web计数的系统,而与语言无关。
We propose a new method of selecting hypotheses for machine transliteration. We generate a set of Chinese, Japanese, and Korean transliteration hypotheses for a given English word. We then use the set of transliteration hypotheses as a guide to finding relevant Web pages and mining contextual information for the transliteration hypotheses from the Web page. Finally, we use the mined information for machine-learning algorithms including support vector machines and maximum entropy model designed to select the correct transliteration hypothesis. In our experiments, our proposed method based on Web mining consistently outperformed systems based on simple Web counts used in previous work, regardless of the language.