Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry

Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry
复制标题

DOI:
10.1016/j.chb.2020.106553
复制
发表时间:
2021-01-01
影响因子:
9.9
通讯作者:
Mossink, Luca D.
Mossink, Luca D.
中科院分区:
心理学1区
文献类型:
--
作者:
Kobis, Nils;Mossink, Luca D.

文献摘要

被引文献

相似文献

公开可用的、强大的自然语言生成算法(NLG)的发布引起了公众的广泛关注和争论。原因之一在于算法据称能够跨不同领域生成类似人类的文本。缺乏使用激励任务来评估人们是否(a)能够区分和(b)是否更喜欢算法生成的文本而不是人类编写的文本的经验证据。我们进行了两项实验,评估对最先进的自然语言生成算法 GPT-2(Ntotal = 830)的行为反应。 GPT-2 使用人类诗歌的相同起始行生成了诗歌样本。从这些样本中,要么随机选择一首诗(人类出环),要么选择最好的一首诗(人类在环),然后与一首人类写的诗进行匹配。在图灵测试的新激励版本中,参与者未能可靠地检测到“人在循环”处理中算法生成的诗歌,但在“人外循环”处理中取得了成功。此外,人们对算法生成的诗歌表现出轻微的厌恶,这与参与者是否了解诗歌的算法起源(透明)或不知晓(不透明)无关。我们讨论这些结果传达了 NLG 算法生成类人文本的性能,并提出了在人类代理实验环境中研究此类学习算法的方法。
The release of openly available, robust natural language generation algorithms (NLG) has spurred much public attention and debate. One reason lies in the algorithms' purported ability to generate humanlike text across various domains. Empirical evidence using incentivized tasks to assess whether people (a) can distinguish and (b) prefer algorithm-generated versus human-written text is lacking. We conducted two experiments assessing behavioral reactions to the state-of-the-art Natural Language Generation algorithm GPT-2 (Ntotal = 830). Using the identical starting lines of human poems, GPT-2 produced samples of poems. From these samples, either a random poem was chosen (Human-out-of-theloop) or the best one was selected (Human-in-the-loop) and in turn matched with a human-written poem. In a new incentivized version of the Turing Test, participants failed to reliably detect the algorithmically generated poems in the Human-in-the-loop treatment, yet succeeded in the Human-out-of-the-loop treatment. Further, people reveal a slight aversion to algorithm-generated poetry, independent on whether participants were informed about the algorithmic origin of the poem (Transparency) or not (Opacity). We discuss what these results convey about the performance of NLG algorithms to produce human-like text and propose methodologies to study such learning algorithms in human-agent experimental settings.