A Labeled Dataset for Investigating Cyberbullying Content Patterns in Instagram

A Labeled Dataset for Investigating Cyberbullying Content Patterns in Instagram
复制标题

DOI:
10.1609/icwsm.v16i1.19376
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Mara Hamlett;G. Powell;Yasin N. Silva;Deborah L. Hall
Mara Hamlett;G. Powell;Yasin N. Silva;Deborah L. Hall
中科院分区:
其他
文献类型:
--
作者:
Mara Hamlett;G. Powell;Yasin N. Silva;Deborah L. Hall

文献摘要

被引文献

相似文献

随着在线交流越来越普遍,网络欺凌事件也变得越来越普遍,尤其是在社交媒体网站上。该领域之前的研究已经研究了网络欺凌的结果、网络欺凌受害/实施的预测因素,以及依赖标记数据集来识别潜在模式的计算检测模型。然而,研究网络欺凌发生时所说内容的工作缺乏,大多数可用的数据集只包括基本标签(网络欺凌与否)。本文提出了一个带注释的Instagram数据集,其中包含关于网络欺凌关键属性的详细标签,例如内容类型、目的、方向性、与其他现象的共现性,以及执行注释的个人的人口统计信息。此外,报告了探索性逻辑回归分析的结果,以说明如何从这个标记的数据集中获得关于网络欺凌及其自动检测的新见解。
As online communication continues to become more prevalent, instances of cyberbullying have also become more common, particularly on social media sites. Previous research in this area has studied cyberbullying outcomes, predictors of cyberbullying victimization/perpetration, and computational detection models that rely on labeled datasets to identify the underlying patterns. However, there is a dearth of work examining the content of what is said when cyberbullying occurs and most of the available datasets include only basic labels (cyberbullying or not). This paper presents an annotated Instagram dataset with detailed labels about key cyberbullying properties, such as the content type, purpose, directionality, and co-occurrence with other phenomena, as well as demographic information about the individuals who performed the annotations. Additionally, results of an exploratory logistic regression analysis are reported to illustrate how new insights about cyberbullying and its automatic detection can be gained from this labeled dataset.