Japanese Pronunciation Prediction as Phrasal Statistical Machine Translation

Japanese Pronunciation Prediction as Phrasal Statistical Machine Translation
复制标题

DOI:
--
复制
发表时间:
2011-11
影响因子:
4
通讯作者:
Jun Hatori;Hisami Suzuki
Jun Hatori;Hisami Suzuki
中科院分区:
材料科学3区
文献类型:
--
作者:
Jun Hatori;Hisami Suzuki

文献摘要

相似文献

本文研究日语文本的发音预测问题。这项任务的难度在于日语字符和单词的发音具有高度的模糊性。以前的方法要么考虑的任务作为一个词级分类问题的基础上的字典,这并没有很好地处理外的词汇(OOV)的话,或只专注于OOV的单词的发音预测,而不考虑文本中的单词发音的上下文消歧。在本文中,我们提出了一个统一的方法短语统计机器翻译(SMT)的框架内,结合了基于字典和基于子串的方法的优势。我们的方法是新颖的,我们结合联合收割机单词和字符为基础的发音从一个字典中的SMT框架:前者捕捉单词发音的特质,而后者提供了灵活性来预测OOV单词的发音。我们表明,基于对各种测试集的广泛评估,我们的模型显着优于以前的最先进的系统,在大多数领域实现了约90%的准确性。
This paper addresses the problem of predicting the pronunciation of Japanese text. The difficulty of this task lies in the high degree of ambiguity in the pronunciation of Japanese characters and words. Previous approaches have either considered the task as a word-level classification problem based on a dictionary, which does not fare well in handling out-of-vocabulary (OOV) words; or solely focused on the pronunciation prediction of OOV words without considering the contextual disambiguation of word pronunciations in text. In this paper, we propose a unified approach within the framework of phrasal statistical machine translation (SMT) that combines the strengths of the dictionary-based and substring-based approaches. Our approach is novel in that we combine wordand character-based pronunciations from a dictionary within an SMT framework: the former captures the idiosyncratic properties of word pronunciation, while the latter provides the flexibility to predict the pronunciation of OOV words. We show that based on an extensive evaluation on various test sets, our model significantly outperforms the previous state-of-the-art systems, achieving around 90% accuracy in most domains.