Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text

Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text
复制标题

DOI:
10.48550/arxiv.2212.12672
复制
发表时间:
2022-12
期刊:
--
影响因子:
--
通讯作者:
Liam Dugan;Daphne Ippolito;Arun Kirubarajan;Sherry Shi;Chris Callison-Burch
Liam Dugan;Daphne Ippolito;Arun Kirubarajan;Sherry Shi;Chris Callison-Burch
中科院分区:
其他
文献类型:
--
作者:
Liam Dugan;Daphne Ippolito;Arun Kirubarajan;Sherry Shi;Chris Callison-Burch

文献摘要

被引文献

相似文献

随着大型语言模型生成的文本激增,理解人类如何与这些文本互动,以及他们是否能够检测到他们正在阅读的文本何时不是由人类作者创作的,变得至关重要。先前关于人工检测生成文本的工作主要集中在整个段落是人工编写或机器生成的情况下。在本文中,我们研究了一个更现实的设置,其中文本从人类书写开始,过渡到由最先进的神经语言模型生成。我们表明,虽然注释者经常在这个任务上挣扎,但注释者的技能有很大的差异,并且给予适当的激励,注释者可以随着时间的推移在这个任务上有所提高。此外,我们进行了详细的比较研究,并分析了各种变量(模型大小、解码策略、微调、提示类型等)如何影响人类的检测性能。最后,我们收集了参与者的错误注释,并使用它们来证明某些文本类型会影响模型产生不同类型的错误,并且某些句子级特征与注释者的选择高度相关。我们发布了RoFT数据集:超过21,000个与错误分类配对的人类注释的集合,以鼓励未来在人类检测和评估生成文本方面的工作。
As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior work on human detection of generated text focuses on the case where an entire passage is either human-written or machine-generated. In this paper, we study a more realistic setting where text begins as human-written and transitions to being generated by state-of-the-art neural language models. We show that, while annotators often struggle at this task, there is substantial variance in annotator skill and that given proper incentives, annotators can improve at this task over time. Furthermore, we conduct a detailed comparison study and analyze how a variety of variables (model size, decoding strategy, fine-tuning, prompt genre, etc.) affect human detection performance. Finally, we collect error annotations from our participants and use them to show that certain textual genres influence models to make different types of errors and that certain sentence-level features correlate highly with annotator selection. We release the RoFT dataset: a collection of over 21,000 human annotations paired with error classifications to encourage future work in human detection and evaluation of generated text.