TOCAB: A Dataset for Chinese Abusive Language Processing
TOCAB: A Dataset for Chinese Abusive Language Processing
复制标题
TOCAB:中文辱骂性语言处理数据集
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Chuan
中科院分区:
文献类型:
--
作者:
I. Chung;Chuan
This paper introduced TOCAB, a larger dataset for Chinese abusive language detection and classification. This dataset contains 121,344 real sentences collected from a social media site. Several baseline systems built by machine learning or deep learning were proposed to test this benchmark. BERT is the best baseline system which achieves F1-scores of 0.886 in detection and 0.781 in classification. The bootstrap aggregating BERT model, a state-of-the-art system, outperforms our BERT baseline system, with F1-scores of 0.893 in detection and 0.782 in classification.