Attribution and Obfuscation of Neural Text Authorship: A Data Mining Perspective

Attribution and Obfuscation of Neural Text Authorship: A Data Mining Perspective
复制标题

DOI:
10.1145/3606274.3606276
复制
发表时间:
2022-10
期刊:
ACM SIGKDD Explorations Newsletter
影响因子:
--
通讯作者:
Adaku Uchendu;Thai Le;Dongwon Lee
Adaku Uchendu;Thai Le;Dongwon Lee
中科院分区:
其他
文献类型:
--
作者:
Adaku Uchendu;Thai Le;Dongwon Lee

文献摘要

相似文献

在隐私研究中,两个相互关联的研究问题日益引起人们的兴趣和重要性,它们是作者归属(AA)和作者身份混淆(AO)。给定人工产物,特别是所讨论的文本t,AA解决方案旨在从许多候选作者中准确地将t归因于其真实作者,而AO解决方案旨在修改t以隐藏其真实作者身份。传统上,作者身份的概念及其随之而来的隐私问题只针对人类作者。然而,近年来,由于神经文本生成(NTG)技术在自然语言处理中的爆炸性进步,能够合成人类质量的开放式文本(所谓的神经文本),现在人们不得不考虑人工、机器或它们的组合的作者。由于神经文本被恶意使用时的含义和潜在威胁,了解传统AA/AO解决方案的局限性并开发新的AA/AO解决方案在处理神经文本方面变得至关重要。因此,本文从数据挖掘的角度对神经文本作者归因和混淆的研究现状进行了综述,并对其局限性和未来的研究方向提出了自己的看法。
Two interlocking research questions of growing interest and importance in privacy research are Authorship Attribution (AA) and Authorship Obfuscation (AO). Given an artifact, especially a text t in question, an AA solution aims to accurately attribute t to its true author out of many candidate authors while an AO solution aims to modify t to hide its true authorship. Traditionally, the notion of authorship and its accompanying privacy concern is only toward human authors. However, in recent years, due to the explosive advancements in Neural Text Generation (NTG) techniques in NLP, capable of synthesizing human-quality openended texts (so-called "neural texts"), one has to now consider authorships by humans, machines, or their combination. Due to the implications and potential threats of neural texts when used maliciously, it has become critical to understand the limitations of traditional AA/AO solutions and develop novel AA/AO solutions in dealing with neural texts. In this survey, therefore, we make a comprehensive review of recent literature on the attribution and obfuscation of neural text authorship from a Data Mining perspective, and share our view on their limitations and promising research directions.