Smaller Language Models are Better Black-box Machine-Generated Text Detectors

Smaller Language Models are Better Black-box Machine-Generated Text Detectors
复制标题

较小的语言模型是更好的黑盒机器生成的文本检测器

DOI:
--
复制
发表时间:
2023
期刊:
arXiv.org
影响因子:
--
通讯作者:
Taylor Berg
Taylor Berg
中科院分区:
--
文献类型:
--
作者:
Fatemehsadat Mireshghallah;Justus Mattern;Sicun Gao;R. Shokri;Taylor Berg

文献摘要

参考文献

被引文献

相似文献

随着流畅的生成语言模型的出现,它可以产生与人类书写的非常相似的令人信服的话语,区分一段文本是机器生成的还是人类书写的变得更加具有挑战性和重要性,因为这些模型可以用来传播错误信息,假新闻,假评论以及模仿某些作者和人物。为此,已经提出了一系列方法来检测机器生成的文本。这些方法中的大多数需要访问目标模型的logit或需要从目标中采样的能力。一种这样的黑盒检测方法依赖于这样的观察:生成的文本在生成器的似然函数下是局部最优的,而人类书写的文本不是。我们发现,总体而言,较小的和部分训练的模型是更好的通用文本检测器:它们可以更精确地检测从小模型和大模型生成的文本。有趣的是,我们发现检测器和生成器是否在相同的数据上训练对检测成功并不重要。例如,OPT-125 M模型在检测ChatGPT代时的AUC为0.81,而来自GPT家族的更大模型GPTJ-6 B的AUC为0.45。
With the advent of fluent generative language models that can produce convincing utterances very similar to those written by humans, distinguishing whether a piece of text is machine-generated or human-written becomes more challenging and more important, as such models could be used to spread misinformation, fake news, fake reviews and to mimic certain authors and figures. To this end, there have been a slew of methods proposed to detect machine-generated text. Most of these methods need access to the logits of the target model or need the ability to sample from the target. One such black-box detection method relies on the observation that generated text is locally optimal under the likelihood function of the generator, while human-written text is not. We find that overall, smaller and partially-trained models are better universal text detectors: they can more precisely detect text generated from both small and larger models. Interestingly, we find that whether the detector and generator were trained on the same data is not critically important to the detection success. For instance the OPT-125M model has an AUC of 0.81 in detecting ChatGPT generations, whereas a larger model from the GPT family, GPTJ-6B, has AUC of 0.45.
DOI: 10.48550/arxiv.2212.12672
发表时间: 2022-12
期刊: --
影响因子: --
作者:
Liam Dugan;Daphne Ippolito;Arun Kirubarajan;Sherry Shi;Chris Callison-Burch
通讯作者: Liam Dugan;Daphne Ippolito;Arun Kirubarajan;Sherry Shi;Chris Callison-Burch