KUISAIL at SemEval-2020 Task 12: BERT-CNN for Offensive Speech Identification in Social Media

KUISAIL at SemEval-2020 Task 12: BERT-CNN for Offensive Speech Identification in Social Media
复制标题

KUISAIL 参加 SemEval-2020 任务 12:社交媒体中攻击性语音识别的 BERT-CNN

DOI:
10.18653/v1/2020.semeval-1.271
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
Deniz Yuret
Deniz Yuret
中科院分区:
--
文献类型:
--
作者:
Ali Safaya;Moutasem Abdullatif;Deniz Yuret

文献摘要

被引文献

相似文献

在本文中,我们描述了我们利用卷积神经网络的预训练BERT模型来完成多语言攻击性语言识别共享任务(OffensEval 2020)的子任务A的方法,该任务是SemEval 2020的一部分。我们表明,将CNN与BERT结合起来比单独使用BERT更好,并且我们强调了在下游任务中使用预训练的语言模型的重要性。我们的系统阿拉伯文宏观平均F1-Score为0.897,排名第4,希腊文为0.843,土耳其文为0.814,排名第3。此外,我们还提供ArabicBERT,这是一组预先训练的阿拉伯语转换语言模型,我们与社区共享。
In this paper, we describe our approach to utilize pre-trained BERT models with Convolutional Neural Networks for sub-task A of the Multilingual Offensive Language Identification shared task (OffensEval 2020), which is a part of the SemEval 2020. We show that combining CNN with BERT is better than using BERT on its own, and we emphasize the importance of utilizing pre-trained language models for downstream tasks. Our system, ranked 4th with macro averaged F1-Score of 0.897 in Arabic, 4th with score of 0.843 in Greek, and 3rd with score of 0.814 in Turkish. Additionally, we present ArabicBERT, a set of pre-trained transformer language models for Arabic that we share with the community.