Phonetic Normalization for Machine Translation of User Generated Content

Phonetic Normalization for Machine Translation of User Generated Content
复制标题

用户生成内容的机器翻译的语音规范化

DOI:
10.18653/v1/d19-5553
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
Guillaume Wisniewski
Guillaume Wisniewski
中科院分区:
--
文献类型:
--
作者:
José Carlos Rosales Núñez;Djamé Seddah;Guillaume Wisniewski

文献摘要

被引文献

相似文献

我们提出了一种方法来纠正嘈杂的用户生成的内容(UGC)在法国的目的是产生一个预处理管道,以改善机器翻译这种非规范语料库。为了做到这一点,我们实现了一个基于字符的神经模型发音器,以产生单词的IPA发音。通过这种方式,我们打算纠正语法,词汇和重音错误经常出现在嘈杂的UGC语料库。我们的方法利用了这样一个事实,即一些错误是由于具有相似发音的单词引起的混淆,这些单词可以使用语音查找表进行纠正,以产生归一化候选人。然后,这些潜在的校正被编码在网格中,并使用语言模型进行排名,以输出最可能的校正短语。与使用其他拼音器相比,我们的方法增强了UGC上基于transformer的机器翻译系统。
We present an approach to correct noisy User Generated Content (UGC) in French aiming to produce a pretreatement pipeline to improve Machine Translation for this kind of non-canonical corpora. In order to do so, we have implemented a character-based neural model phonetizer to produce IPA pronunciations of words. In this way, we intend to correct grammar, vocabulary and accentuation errors often present in noisy UGC corpora. Our method leverages on the fact that some errors are due to confusion induced by words with similar pronunciation which can be corrected using a phonetic look-up table to produce normalization candidates. These potential corrections are then encoded in a lattice and ranked using a language model to output the most probable corrected phrase. Compare to using other phonetizers, our method boosts a transformer-based machine translation system on UGC.