Collapsed Consonant and Vowel Models: New Approaches for English-Persian Transliteration and Back-Transliteration
Collapsed Consonant and Vowel Models: New Approaches for English-Persian Transliteration and Back-Transliteration
复制标题
折叠的辅音和元音模型:英波斯音译和回译的新方法
DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
A. Turpin
中科院分区:
文献类型:
--
作者:
Sarvnaz Karimi;Falk Scholer;A. Turpin
Most current machine transliteration systems employ a corpus of known sourcetarget word pairs to train their system, and typically evaluate their systems on a similar corpus. In this paper we explore the performance of transliteration systems on corpora that are varied in a controlled way. In particular, we control the number, and prior language knowledge of human transliterators used to construct the corpora, and the origin of the source words that make up the corpora. We find that the word accuracy of automated transliteration systems can vary by up to 30% (in absolute terms) depending on the corpus on which they are run. We conclude that at least four human transliterators should be used to construct corpora for evaluating automated transliteration systems; and that although absolute word accuracy metrics may not translate across corpora, the relative rankings of system performance remains stable across differing corpora.