Capsule Network Over Pre-Trained Language Model and User Writing Styles for Authorship Attribution on Short Texts

Capsule Network Over Pre-Trained Language Model and User Writing Styles for Authorship Attribution on Short Texts
复制标题

DOI:
10.1145/3562007.3562027
复制
发表时间:
2022-08
期刊:
Proceedings of the 2022 3rd International Conference on Control, Robotics and Intelligent System
影响因子:
--
通讯作者:
Zeping Huang;M. Iwaihara
Zeping Huang;M. Iwaihara
中科院分区:
其他
文献类型:
--
作者:
Zeping Huang;M. Iwaihara

文献摘要

相似文献

作者归属(Authorship Attribution,AA)是作者分析和文本分类的一个子领域,它将文本归因于一组封闭的潜在作者中的正确作者。由于短文本通常包含的作者信息较少,因此对短文本的作者身份归因往往比对长文本的作者身份归因更具挑战性。近年来,预训练语言模型的广泛使用极大地提高了文本分类任务的准确率。在本文中,我们提出了一种使用带有胶囊网络的预训练语言模型BERTweet的模型来解决推文的作者归属问题。BERTweet是第一个针对英语推文的大规模特定领域的预训练语言模型,可以生成高质量的推文句子表示。我们将BERTweet与胶囊网络相结合,这在捕捉句子表征的深层特征方面特别强大。因此,BERTweet和Capsule都帮助我们在AA任务上取得了显著的改进。我们还将用户写作风格融入到我们的模型中。我们设计了新的胶囊网络体系结构,将多个胶囊层结合在一起,从推文和用户写作风格生成表示,提高了预测精度和稳健性。我们的实验结果表明,我们的BERTweet_Capsage_UWS组合在已知的tweet AA数据集上显示了最先进的结果。
Authorship Attribution (AA) is a sub-field of Authorship Analysis and text classification, attributing a text to the correct author among a closed set of potential authors. Since short texts usually contain less information about the author, authorship attribution on short texts is often more challenging than authorship attribution on long texts. Recently, the widespread use of pre-trained language models has greatly improved the accuracy of text classification tasks. In this paper, we propose a model which uses the pre-trained language model BERTweet with capsule networks, to solve the authorship attribution on tweets. BERTweet is the first large-scale domain-specific pre-trained language model for English tweets, which can generate high-quality sentence representations of tweets. We combine BERTweet with capsule networks which are particularly powerful at capturing deep features of sentence representations. Thus, both BERTweet and capsule help us achieve remarkable improvements on AA tasks. We also incorporate user writing styles into our model. We design new architectures of capsule networks which combine multiple capsule layers, for generating representations from tweets and user writing styles, improving prediction accuracy and robustness. Our experimental results show that our BERTweet_Capsule_UWS combination shows the state-of-the-art result on the known tweet AA dataset.