Semi-supervised Learning for Cyberbullying Detection in Social Networks

Semi-supervised Learning for Cyberbullying Detection in Social Networks
复制标题

社交网络中网络欺凌检测的半监督学习

DOI:
--
复制
发表时间:
2014
期刊:
Australasian Database Conference
影响因子:
--
通讯作者:
C. Pang
C. Pang
中科院分区:
--
文献类型:
--
作者:
V. Nahar;S. Al;Xue Li;C. Pang

文献摘要

被引文献

相似文献

目前的网络欺凌检测方法大多是静态的:它们无法有效地处理嘈杂、不平衡或流数据。现有的网络欺凌检测研究主要是监督学习方法,假设数据有足够的预标记。然而,在流数据中只有少量标签可用的现实情况下,这是不切实际的。在本文中,我们提出了一种半监督学习方法,该方法将增加训练数据样本并应用模糊支持向量机算法。增强训练技术自动从未标记的流文本中提取和扩大训练集,同时利用提供的非常小的训练集作为初始输入进行学习。实验结果表明,所提出的增强方法优于所有其他方法,并且适用于没有足够标记实例可用于训练的现实情况。对于所提出的模糊支持向量机方法,我们处理由流文本生成的复杂和多维数据,其中特征的重要性被区分为决策函数。在不同的实验场景下进行的评价表明,所提出的模糊支持向量机相对于所有其他方法具有优越性。
Current approaches on cyberbullying detection are mostly static: they are unable to handle noisy, imbalanced or streaming data efficiently. Existing studies on cyberbullying detection are mainly supervised learning approaches, assuming data is sufficiently pre-labelled. However this is impractical in the real-world situation where only a small number of labels are available in streaming data. In this paper, we propose a semi-supervised leaning approach that will augment training data samples and apply a fuzzy SVM algorithm. The augmented training technique automatically extracts and enlarges training set from the unlabelled streaming text, while learning is conducted by utilising a very small training set provided as an initial input. The experimental results indicate that the proposed augmented approach outperformed all other methods, and is suitable in the real-world situations, where sufficiently labelled instances are not available for training. For the proposed fuzzy SVM approach we handle complex and multidimensional data generated by streaming text, where the importance of features are discriminated for the decision function. The evaluation conducted on different experimental scenarios indicates the superiority of the proposed fuzzy SVM against all other methods.