A Fuzzy Approach to Text Classification With Two-Stage Training for Ambiguous Instances

A Fuzzy Approach to Text Classification With Two-Stage Training for Ambiguous Instances
复制标题

DOI:
10.1109/tcss.2019.2892037
复制
发表时间:
2019-03
影响因子:
5
通讯作者:
Han Liu;P. Burnap;Wafa Alorainy;M. Williams
Han Liu;P. Burnap;Wafa Alorainy;M. Williams
中科院分区:
计算机科学2区
文献类型:
--
作者:
Han Liu;P. Burnap;Wafa Alorainy;M. Williams

文献摘要

被引文献

相似文献

情感分析是文本挖掘和机器学习的一个非常热门的应用领域。流行的方法包括支持向量机,朴素贝叶斯,决策树和深度神经网络。然而,这些方法通常属于判别式学习,其目的是在地面真实存在的情况下,以明确的结果区分一个类与其他类。在文本分类的上下文中,实例自然是模糊的(在某些应用领域中可以是多标签的),因此不被认为是明确的,特别是考虑到为文本中的情感分配的标签代表多个人类注释者的主观意见的一致水平,而不是无可争辩的事实。这促使研究人员开发模糊方法,这种方法通常通过生成学习来训练分类器,即,模糊分类器用于测量实例属于每个类的程度。传统的模糊方法通常涉及单个模糊分类器的生成,并采用固定的解模糊规则,输出具有最大隶属度的类。使用具有上述固定的去模糊化规则的单个模糊分类器可能会使分类器遇到情感数据上的文本歧义情况,即,实例可以获得对正类和负类两者的相等隶属度。在本文中,我们重点关注网络仇恨分类,因为通过社交媒体传播仇恨言论可能会对社会凝聚力产生破坏性影响,并导致地区和社区紧张局势。因此,自动检测网络仇恨已成为一个优先研究领域。特别是,我们提出了一种改进的模糊方法与两阶段的训练,用于处理文本歧义和分类四种类型的仇恨言论,即宗教,种族,残疾和性取向,并将其性能与这些流行的方法以及一些现有的模糊方法进行比较,而特征是通过词袋和词嵌入特征提取方法以及相关性来准备的,的特征子集选择方法。实验结果表明,在大多数情况下,所提出的模糊方法优于其他方法。
Sentiment analysis is a very popular application area of text mining and machine learning. The popular methods include support vector machine, naive bayes, decision trees, and deep neural networks. However, these methods generally belong to discriminative learning, which aims to distinguish one class from others with a clear-cut outcome, under the presence of ground truth. In the context of text classification, instances are naturally fuzzy (can be multilabeled in some application areas) and thus are not considered clear-cut, especially given the fact that labels assigned to sentiment in text represent an agreed level of subjective opinion for multiple human annotators rather than indisputable ground truth. This has motivated researchers to develop fuzzy methods, which typically train classifiers through generative learning, i.e., a fuzzy classifier is used to measure the degree to which an instance belongs to each class. Traditional fuzzy methods typically involve generation of a single fuzzy classifier and employ a fixed rule of defuzzification outputting the class with the maximum membership degree. The use of a single fuzzy classifier with the above-fixed rule of defuzzification is likely to get the classifier encountering the text ambiguity situation on sentiment data, i.e., an instance may obtain equal membership degrees to both the positive and negative classes. In this paper, we focus on cyberhate classification, since the spread of hate speech via social media can have disruptive impacts on social cohesion and lead to regional and community tensions. Automatic detection of cyberhate has thus become a priority research area. In particular, we propose a modified fuzzy approach with two-stage training for dealing with text ambiguity and classifying four types of hate speech, namely, religion, race, disability, and sexual orientation—and compare its performance with those popular methods as well as some existing fuzzy approaches, while the features are prepared through the bag-of-words and word embedding feature extraction methods alongside the correlation-based feature subset selection method. The experimental results show that the proposed fuzzy method outperforms the other methods in most cases.