A Human Judgement Corpus and a Metric for Arabic MT Evaluation

A Human Judgement Corpus and a Metric for Arabic MT Evaluation
复制标题

人类判断语料库和阿拉伯语机器翻译评估指标

DOI:
10.3115/v1/d14-1026
复制
发表时间:
2014
期刊:
ArXiv
影响因子:
--
通讯作者:
Kemal Oflazer
Kemal Oflazer
中科院分区:
--
文献类型:
--
作者:
Houda Bouamor;Hanan Alshikhabobakr;Behrang Mohit;Kemal Oflazer

文献摘要

被引文献

相似文献

我们提出了一个人类判断数据集和一个适用于阿拉伯语机器翻译评估的度量标准。我们的MeumscaleDataSet是第一个具有高注释质量的Arabi数据集。我们使用该数据集将BLEU分数调整为阿拉伯语。我们的评分(AL-BLEU)为假设和参照词的词干和形态匹配提供了部分学分。我们在我们的人类判断语料库上对BLEU、流星和BLEU进行了评估,结果表明,AL-BLEU与人类判断的相关性最高。我们正在向研究界发布数据集和软件。
We present a human judgments datasetand an adapted metric for evaluation ofArabic machine translation. Our mediumscaledataset is the first of its kind for Arabicwith high annotation quality. We usethe dataset to adapt the BLEU score forArabic. Our score (AL-BLEU) providespartial credits for stem and morphologicalmatchings of hypothesis and referencewords. We evaluate BLEU, METEOR andAL-BLEU on our human judgments corpusand show that AL-BLEU has the highestcorrelation with human judgments. Weare releasing the dataset and software tothe research community.