The Multilingual Amazon Reviews Corpus

The Multilingual Amazon Reviews Corpus
复制标题

DOI:
10.18653/v1/2020.emnlp-main.369
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Phillip Keung;Y. Lu;György Szarvas;Noah A. Smith
Phillip Keung;Y. Lu;György Szarvas;Noah A. Smith
中科院分区:
其他
文献类型:
--
作者:
Phillip Keung;Y. Lu;György Szarvas;Noah A. Smith

文献摘要

被引文献

相似文献

我们提出了多语言亚马逊评论语料库(MARC),一个大规模的亚马逊评论多语言文本分类的集合。该语料库包含英语,日语,德语,法语,西班牙语和中文的评论,这些评论是在2015年至2019年期间收集的。数据集中的每个记录包含评论文本、评论标题、星星评级、匿名评论者ID、匿名产品ID和粗粒度产品类别(例如,“书籍”、“器具”等)语料库在5个可能的星星评级中保持平衡,因此每个评级占每种语言评论的20%。对于每种语言,在训练集、开发集和测试集中分别有200,000、5,000和5,000条评论。我们通过对评论数据的多语言BERT模型进行微调,报告了监督文本分类和零次跨语言迁移学习的基线结果。我们建议使用平均绝对误差(MAE),而不是分类准确度为这项任务,因为MAE帐户的顺序性质的评级。
We present the Multilingual Amazon Reviews Corpus (MARC), a large-scale collection of Amazon reviews for multilingual text classification. The corpus contains reviews in English, Japanese, German, French, Spanish, and Chinese, which were collected between 2015 and 2019. Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID, and the coarse-grained product category (e.g., 'books', 'appliances', etc.) The corpus is balanced across the 5 possible star ratings, so each rating constitutes 20% of the reviews in each language. For each language, there are 200,000, 5,000, and 5,000 reviews in the training, development, and test sets, respectively. We report baseline results for supervised text classification and zero-shot cross-lingual transfer learning by fine-tuning a multilingual BERT model on reviews data. We propose the use of mean absolute error (MAE) instead of classification accuracy for this task, since MAE accounts for the ordinal nature of the ratings.