TOCAB: A Dataset for Chinese Abusive Language Processing

TOCAB: A Dataset for Chinese Abusive Language Processing
复制标题

TOCAB:中文辱骂性语言处理数据集

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Information Reuse and Integration
影响因子:
--
通讯作者:
Chuan
Chuan
中科院分区:
--
文献类型:
--
作者:
I. Chung;Chuan

文献摘要

被引文献

相似文献

本文介绍了一种用于汉语滥用语言检测和分类的大型数据集TOCAB。该数据集包含从社交媒体网站收集的121,344个真实句子。提出了几个通过机器学习或深度学习构建的基线系统来测试这个基准。BERT是最好的基线系统,检测得分为0.886,分类得分为0.781。bootstrap aggregating BERT模型是最先进的系统,优于我们的BERT基线系统,在检测方面的f1得分为0.893,在分类方面的f1得分为0.782。
This paper introduced TOCAB, a larger dataset for Chinese abusive language detection and classification. This dataset contains 121,344 real sentences collected from a social media site. Several baseline systems built by machine learning or deep learning were proposed to test this benchmark. BERT is the best baseline system which achieves F1-scores of 0.886 in detection and 0.781 in classification. The bootstrap aggregating BERT model, a state-of-the-art system, outperforms our BERT baseline system, with F1-scores of 0.893 in detection and 0.782 in classification.