S3 map: Semisupervised aspect-based sentiment analysis with masked aspect prediction

S3 map: Semisupervised aspect-based sentiment analysis with masked aspect prediction
复制标题

DOI:
10.1016/j.knosys.2023.110513
复制
发表时间:
2023-03
影响因子:
8.8
通讯作者:
Zhiyao Yang;Bing Wang;Ximing Li;Wenting Wang;Jihong Ouyang
Zhiyao Yang;Bing Wang;Ximing Li;Wenting Wang;Jihong Ouyang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhiyao Yang;Bing Wang;Ximing Li;Wenting Wang;Jihong Ouyang

文献摘要

相似文献

基于方面的情感分析(ABSA)是一种在方面层面检测句子情感极性的细粒度任务。为了解决这个问题,ABSA训练样本必须用方面词和相应的情感极性进行标注。然而,收集这种细粒度的训练样本既昂贵又耗时。因此,可用的ABSA训练样本往往很少。为了打破ABSA的数据稀缺性挑战,我们研究了半监督基于方面的情感分析(SemiABSA),它使用有限数量的昂贵标记句子和更多未标记但更便宜的句子来训练ABSA模型。我们提出了一种基于自我训练范式的semi- absa框架,即半监督基于方面的情感分析与屏蔽方面预测(s3 map)。我们对未标注的句子形成伪方面词和伪情感极性,并改进模型训练。具体来说,基于bert编码器的掩码方面预测(mask aspect prediction, MAP)任务实现了伪方面词的生成。基于s3图谱,我们从多个角度深入研究了SemiABSA的潜力。实证结果表明,s3 map可以通过利用未标记的句子来持续提高性能,即使这些句子来自不同的领域。
Aspect-based sentiment analysis (ABSA) refers to a fine-grained task of detecting the sentiment polarities of sentences at the aspect level. To resolve this task, ABSA training samples must be annotated with aspect words and the corresponding sentiment polarities. However, collecting such fine-grained training samples is expensive and time-consuming. Therefore the available ABSA training samples are often scarce. To break the data scarcity challenge of ABSA, we investigate semi-supervised aspect-based sentiment analysis (SemiABSA), which trains ABSA models using a limited amount of expensive labeled sentences and more unlabeled-yet-cheaper sentences. We propose a novel SemiABSA framework, namely semi-supervised aspect-based sentiment analysis with masked aspect prediction (S 3 map), built on the self-training paradigm. We form pseudo-aspect words and pseudo-sentiment polarities for unlabeled sentences and improve model training. Specifically, a BERT-encoder-based masked aspect prediction (MAP) task achieves the pseudo-aspect words generation. Based on S 3 map, we thoroughly investigate the potential of SemiABSA from various perspectives. The empirical results show that S 3 map can consistently improve performance by leveraging unlabeled sentences, even those from different domains.