WebCrawl African : A Multilingual Parallel Corpora for African Languages

WebCrawl African : A Multilingual Parallel Corpora for African Languages
复制标题

WebCrawl African:非洲语言多语言并行语料库

DOI:
--
复制
发表时间:
2022
期刊:
Conference on Machine Translation
影响因子:
--
通讯作者:
C. Viswanathan
C. Viswanathan
中科院分区:
--
文献类型:
--
作者:
Pavanpankaj Vegi;J. Sivabhavani;Biswajit Paul;Abhinav Mishra;Prashant Banjare;Krtin Kumar;C. Viswanathan

文献摘要

参考文献

被引文献

相似文献

WebCrawl African是由人工智能和机器人实验室中心的ANVITA机器翻译团队编译的非洲语言池的混合域多语言并行语料库,主要用于加速低资源和极低资源机器翻译的研究,并且是提交给WMT 2022的数据轨道下非洲语言大规模机器翻译评估共享任务的一部分。该语料库是通过网络数据挖掘编译的,包括69.5万个平行句子,涵盖英语和15种非洲语言的74种不同语言对,其中许多属于低和极低资源类别。作为语料库有用性的衡量标准,通过将WebCrawl非洲语料库与现有语料库相结合来训练24种非洲语言到英语的MNMT模型,并且在FLORES 200上的评估显示,包含WebCrawl非洲语料库可以将15个非洲到英语翻译方向中的12个的BLEU得分提高0.01-1.66,甚至提高0.18- 0.68对于9个非洲到英语的翻译方向中的4个,它们不是WebCrawl非洲语料库的一部分。与OPUS公共资源库相比,WebCrawl非洲语料库包含了更多的语言对的平行句子。本数据描述文件记录了语料库的创建和沿着的结果以及数据表。WebCrawl非洲语料库托管在github存储库上。
WebCrawl African is a mixed domain multilingual parallel corpora for a pool of African languages compiled by ANVITA machine translation team of Centre for Artificial Intelligence and Robotics Lab, primarily for accelerating research on low-resource and extremely low-resource machine translation and is part of the submission to WMT 2022 shared task on Large-Scale Machine Translation Evaluation for African Languages under the data track. The corpora is compiled through web data mining and comprises 695K parallel sentences spanning 74 different language pairs from English and 15 African languages, many of which fall under low and extremely low resource categories. As a measure of corpora usefulness, a MNMT model for 24 African languages to English is trained by combining WebCrawl African corpora with existing corpus and evaluation on FLORES200 shows that inclusion of WebCrawl African corpora could improve BLEU score by 0.01-1.66 for 12 out of 15 African to English translation directions and even by 0.18-0.68 for the 4 out of 9 African to English translation directions which are not part of WebCrawl African corpora. WebCrawl African corpora includes more parallel sentences for many language pairs in comparison to OPUS public repository. This data description paper captures creation of corpora and results obtained along with datasheets. The WebCrawl African corpora is hosted on github repository.
WMT™22 非洲语言大规模机器翻译评估共享任务的调查结果
DOI: --
发表时间: 2022
期刊: Association for Computational Linguistics
影响因子: --
作者:
Adelani, David;Ibn Alam, Md Mahfuz;Anastasopoulos, Antonios;Bhagia, Akshita;Costa-jussà, Marta R.;Dodge, Jesse;Faisal, Fahim;Federmann, Christian;Fedorova, Natalia;Guzmán, Francisco
通讯作者: Guzmán, Francisco