Multilingual Transliteration Using Feature based Phonetic Method

Multilingual Transliteration Using Feature based Phonetic Method
复制标题

使用基于特征的语音方法进行多语言音译

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
R. Sproat
R. Sproat
中科院分区:
--
文献类型:
--
作者:
Su;Kyoung;R. Sproat

文献摘要

被引文献

相似文献

本文研究了基于语音评分法的命名实体音译。语音方法是利用语音特征和精心设计的伪特征进行计算。使用可比语料库,用四种语言(阿拉伯语、汉语、印地语和朝鲜语)和一种源语言(英语)对提议的方法进行了测试。该方法是在Tao et al.(2006)最初提出的语音方法的基础上发展而来的。与Tao等人(2006)基于纯语言知识构建的语音方法不同,本研究的方法是使用Winnow机器学习算法进行训练的。与之前的研究相比,印地语和阿拉伯语有了显著的进步。此外,我们还证明了该方法在与目标语言不同的语言数据上进行训练时也可以获得可比较的结果。该方法既可以用最少的数据,也可以在没有目标语言数据的情况下应用于各种语言。
In this paper we investigate named entity transliteration based on a phonetic scoring method. The phonetic method is computed using phonetic features and carefully designed pseudo features. The proposed method is tested with four languages – Arabic, Chinese, Hindi and Korean – and one source language – English, using comparable corpora. The proposed method is developed from the phonetic method originally proposed in Tao et al. (2006). In contrast to the phonetic method in Tao et al. (2006) constructed on the basis of pure linguistic knowledge, the method in this study is trained using the Winnow machine learning algorithm. There is salient improvement in Hindi and Arabic compared to the previous study. Moreover, we demonstrate that the method can also achieve comparable results, when it is trained on language data different from the target language. The method can be applied both with minimal data, and without target language data for various languages.