Linguistic-Relationships-Based Approach for Improving Word Alignment

Linguistic-Relationships-Based Approach for Improving Word Alignment
复制标题

DOI:
10.1145/3133323
复制
发表时间:
2017-10
期刊:
ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)
影响因子:
--
通讯作者:
Phuoc Tran;Dinh Dien;Tan Le;Long H. B. Nguyen
Phuoc Tran;Dinh Dien;Tan Le;Long H. B. Nguyen
中科院分区:
其他
文献类型:
--
作者:
Phuoc Tran;Dinh Dien;Tan Le;Long H. B. Nguyen

文献摘要

被引文献

相似文献

无监督词对齐(如吉萨++)广泛应用于基于短语的统计机器翻译中。模型的质量与双语语料库的规模和质量成正比。然而,对于低资源的语言对,如汉语和越南语,无监督的词对齐的结果有时是低质量的,由于稀疏的数据。此外,该模型没有利用语言关系来提高词对齐的性能。汉语和越南语有着相同的语言类型和密切的语言关系。在本文中,我们将语言关系的特点融入到词对齐模型中,以提高汉越词对齐的质量。这些语言关系是汉越语和实词。实验结果表明,该方法提高了词对齐的性能以及机器翻译的质量。
The unsupervised word alignments (such as GIZA++) are widely used in the phrase-based statistical machine translation. The quality of the model is proportional to the size and the quality of the bilingual corpus. However, for low-resource language pairs such as Chinese and Vietnamese, a result of unsupervised word alignment sometimes is of low quality due to the sparse data. In addition, this model does not take advantage of the linguistic relationships to improve performance of word alignment. Chinese and Vietnamese have the same language type and have close linguistic relationships. In this article, we integrate the characteristics of linguistic relationships into the word alignment model to enhance the quality of Chinese-Vietnamese word alignment. These linguistic relationships are Sino-Vietnamese and content word. The experimental results showed that our method improved the performance of word alignment as well as the quality of machine translation.