Practical Comparable Data Collection for Low-Resource Languages via Images
Practical Comparable Data Collection for Low-Resource Languages via Images
复制标题
DOI:
--
复制
发表时间:
2020-04
期刊:
影响因子:
--
通讯作者:
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
中科院分区:
文献类型:
--
作者:
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
We propose a method of curating high-quality comparable training data for low-resource languages with monolingual annotators. Our method involves using a carefully selected set of images as a pivot between the source and target languages by getting captions for such images in both languages independently. Human evaluations on the English-Hindi comparable corpora created with our method show that 81.1% of the pairs are acceptable translations, and only 2.47% of the pairs are not translations at all. We further establish the potential of the dataset collected through our approach by experimenting on two downstream tasks - machine translation and dictionary extraction. All code and data are available at this https URL.