UPTON: Preventing Authorship Leakage from Public Text Release via Data Poisoning

UPTON: Preventing Authorship Leakage from Public Text Release via Data Poisoning
复制标题

DOI:
10.18653/v1/2023.findings-emnlp.800
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Ziyao Wang;Thai Le;Dongwon Lee
Ziyao Wang;Thai Le;Dongwon Lee
中科院分区:
其他
文献类型:
--
作者:
Ziyao Wang;Thai Le;Dongwon Lee

文献摘要

相似文献

考虑这样一种情况:拥有许多公共著作的作者(例如活动家、举报人)希望“匿名”写作,而攻击者可能已经基于包括作者在内的公共著作建立了作者归属(AA)模型。为了实现她的愿望,我们提出了一个问题“能否让公开发布的作品 T 变得不可归属,从而使在 T 上训练的 AA 模型无法很好地归因其作者身份?”针对这个问题,我们提出了一种新颖的解决方案 UPTON,它利用黑盒数据中毒方法来削弱训练样本中的作者特征,使发布的文本变得不可学习。它与以前的混淆工作不同,例如修改测试样本的对抗性攻击或仅在触发词发生时改变模型输出的后门工作。使用四个作者数据集(IMDb10、IMDb64、Enron 和 WJO),我们进行了实证验证,其中 UPTON 成功地将 AA 模型的准确性降低到不切实际的水平(~35%),同时保持文本仍然可读(语义相似度>0.9)。 UPTON 对于已经接受过作者可用的干净写作训练的 AA 模型仍然有效。
Consider a scenario where an author-e.g., activist, whistle-blower, with many public writings wishes to write"anonymously"when attackers may have already built an authorship attribution (AA) model based off of public writings including those of the author. To enable her wish, we ask a question"Can one make the publicly released writings, T, unattributable so that AA models trained on T cannot attribute its authorship well?"Toward this question, we present a novel solution, UPTON, that exploits black-box data poisoning methods to weaken the authorship features in training samples and make released texts unlearnable. It is different from previous obfuscation works-e.g., adversarial attacks that modify test samples or backdoor works that only change the model outputs when triggering words occur. Using four authorship datasets (IMDb10, IMDb64, Enron, and WJO), we present empirical validation where UPTON successfully downgrades the accuracy of AA models to the impractical level (~35%) while keeping texts still readable (semantic similarity>0.9). UPTON remains effective to AA models that are already trained on available clean writings of authors.