Investigating the Effect of Machine-Translation on Automated Classification of Toxic Comments

Investigating the Effect of Machine-Translation on Automated Classification of Toxic Comments
复制标题

DOI:
10.1109/mass56207.2022.00120
复制
发表时间:
2022-10
期刊:
2022 IEEE 19th International Conference on Mobile Ad Hoc and Smart Systems (MASS)
影响因子:
--
通讯作者:
J. Roy;Siddhi Suresh;Mohamed ElSayed;Ronie Rocca;Ziqian Dong;Huanying Gu;N. S. Artan
J. Roy;Siddhi Suresh;Mohamed ElSayed;Ronie Rocca;Ziqian Dong;Huanying Gu;N. S. Artan
中科院分区:
其他
文献类型:
--
作者:
J. Roy;Siddhi Suresh;Mohamed ElSayed;Ronie Rocca;Ziqian Dong;Huanying Gu;N. S. Artan

文献摘要

相似文献

本文讨论了机器翻译后自动有毒评论分类性能的研究结果。我们首先针对五种语言的非英语维基百科讨论页面的评论测试了 Google Perspective API,然后测试了它们的英文翻译(使用 Google 的 Cloud Translate API 生成)。除了提供 Perspective 在五种语言中当前性能的基准之外,还可以对机器翻译如何改变分类进行比较。我们表明,翻译前和翻译后分类之间的分歧程度在很大程度上取决于所使用的语言。这些评论来自 Kaggle 数据集,我们对其进行过滤,以确保使用简单标点符号的单语评论。结果显示,超过 84% 的法语、意大利语和西班牙语评论在翻译前和翻译后获得相同等级,而葡萄牙语和俄语在测试的五种语言中表现最差,F 分数低于 0.6。
This paper discusses the research findings on the performance of automated toxic comment classification following machine translation. We tested Google Perspective API first on comments from non-English Wikipedia talk pages in five languages, and then on their English translation (generated with Google's Cloud Translate API). In addition to giving baselines on the current performance of Perspective in five languages, this allows for comparison on how machine-translation alters the classification. We show that the level of disagreement between pre- and post-translation classification is heavily dependent on the language used. The comments come from a Kaggle dataset and we filter them to ensure monolingual comments with simple punctuation. Results show above 84% of the French, Italian and Spanish comments received the same class pre- and post-translation, while Portuguese and Russian performed the worst of the five languages tested, with F-scores below 0.6.