A Human Judgement Corpus and a Metric for Arabic MT Evaluation
A Human Judgement Corpus and a Metric for Arabic MT Evaluation
复制标题
人类判断语料库和阿拉伯语机器翻译评估指标
DOI:
10.3115/v1/d14-1026
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Kemal Oflazer
中科院分区:
文献类型:
--
作者:
Houda Bouamor;Hanan Alshikhabobakr;Behrang Mohit;Kemal Oflazer
We present a human judgments datasetand an adapted metric for evaluation ofArabic machine translation. Our mediumscaledataset is the first of its kind for Arabicwith high annotation quality. We usethe dataset to adapt the BLEU score forArabic. Our score (AL-BLEU) providespartial credits for stem and morphologicalmatchings of hypothesis and referencewords. We evaluate BLEU, METEOR andAL-BLEU on our human judgments corpusand show that AL-BLEU has the highestcorrelation with human judgments. Weare releasing the dataset and software tothe research community.