Collapsed Consonant and Vowel Models: New Approaches for English-Persian Transliteration and Back-Transliteration

Collapsed Consonant and Vowel Models: New Approaches for English-Persian Transliteration and Back-Transliteration
复制标题

折叠的辅音和元音模型:英波斯音译和回译的新方法

DOI:
--
复制
发表时间:
2007
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
A. Turpin
A. Turpin
中科院分区:
--
文献类型:
--
作者:
Sarvnaz Karimi;Falk Scholer;A. Turpin

文献摘要

被引文献

相似文献

大多数当前的机器音译系统使用已知源目标词对的语料库来训练他们的系统,并且通常在类似的语料库上评估他们的系统。在本文中,我们探讨了音译系统在受控变化的语料库上的性能。特别是,我们控制了用于构建语料库的人类音译者的数量和先验语言知识,以及构成语料库的源词的来源。我们发现,自动音译系统的单词准确性根据其运行的语料库可以变化高达30%(以绝对值计算)。我们的结论是,至少应该使用四名人类音译员来构建用于评估自动音译系统的语料库;而且,尽管绝对的单词准确性指标可能无法在不同的语料库中转换,但系统性能的相对排名在不同的语料库中保持稳定。
Most current machine transliteration systems employ a corpus of known sourcetarget word pairs to train their system, and typically evaluate their systems on a similar corpus. In this paper we explore the performance of transliteration systems on corpora that are varied in a controlled way. In particular, we control the number, and prior language knowledge of human transliterators used to construct the corpora, and the origin of the source words that make up the corpora. We find that the word accuracy of automated transliteration systems can vary by up to 30% (in absolute terms) depending on the corpus on which they are run. We conclude that at least four human transliterators should be used to construct corpora for evaluating automated transliteration systems; and that although absolute word accuracy metrics may not translate across corpora, the relative rankings of system performance remains stable across differing corpora.