Adversarial Semantic Collisions

Adversarial Semantic Collisions
复制标题

DOI:
10.18653/v1/2020.emnlp-main.344
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Congzheng Song;Alexander M. Rush;Vitaly Shmatikov
Congzheng Song;Alexander M. Rush;Vitaly Shmatikov
中科院分区:
其他
文献类型:
--
作者:
Congzheng Song;Alexander M. Rush;Vitaly Shmatikov

文献摘要

被引文献

相似文献

我们研究语义冲突:语义上不相关但被NLP模型判断为相似的文本。我们开发了基于梯度的方法来生成语义冲突,并证明了许多任务的最先进的模型,这些任务依赖于分析文本的含义和相似性-包括释义识别,文档检索,响应建议和提取摘要-容易受到语义冲突的影响。例如,给定一个目标查询,将特制的冲突插入到不相关的文档中可以将其检索排名从1000移到前3。我们将展示如何生成语义冲突,逃避基于困惑的过滤,并讨论其他潜在的缓解。我们的代码可在https://github.com/csong27/collision-bert上获得。
We study semantic collisions: texts that are semantically unrelated but judged as similar by NLP models. We develop gradient-based approaches for generating semantic collisions and demonstrate that state-of-the-art models for many tasks which rely on analyzing the meaning and similarity of texts-- including paraphrase identification, document retrieval, response suggestion, and extractive summarization-- are vulnerable to semantic collisions. For example, given a target query, inserting a crafted collision into an irrelevant document can shift its retrieval rank from 1000 to top 3. We show how to generate semantic collisions that evade perplexity-based filtering and discuss other potential mitigations. Our code is available at https://github.com/csong27/collision-bert.