Practical Comparable Data Collection for Low-Resource Languages via Images

Practical Comparable Data Collection for Low-Resource Languages via Images
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
中科院分区:
其他
文献类型:
--
作者:
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig

文献摘要

相似文献

我们提出了一种为低资源语言和单语注释者管理高质量可比训练数据的方法。我们的方法包括使用一组精心挑选的图像作为源语言和目标语言之间的枢轴,通过独立地获得这两种语言的图像标题。用我们的方法创建的英语-印地语可比语料库上的人工评估显示,81.1%的对是可接受的翻译,只有2.47%的对根本不是翻译。我们通过对两个下游任务-机器翻译和词典提取-进行实验,进一步确定了通过我们的方法收集的数据集的潜力。所有代码和数据都可以在这个https URL上找到。
We propose a method of curating high-quality comparable training data for low-resource languages with monolingual annotators. Our method involves using a carefully selected set of images as a pivot between the source and target languages by getting captions for such images in both languages independently. Human evaluations on the English-Hindi comparable corpora created with our method show that 81.1% of the pairs are acceptable translations, and only 2.47% of the pairs are not translations at all. We further establish the potential of the dataset collected through our approach by experimenting on two downstream tasks - machine translation and dictionary extraction. All code and data are available at this https URL.