TruthfulQA: Measuring How Models Mimic Human Falsehoods

TruthfulQA: Measuring How Models Mimic Human Falsehoods
复制标题

DOI:
10.18653/v1/2022.acl-long.229
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
Stephanie C. Lin;Jacob Hilton;Owain Evans
Stephanie C. Lin;Jacob Hilton;Owain Evans
中科院分区:
其他
文献类型:
--
作者:
Stephanie C. Lin;Jacob Hilton;Owain Evans

文献摘要

被引文献

相似文献

我们提出了一个基准,用于衡量语言模型在回答问题时是否真实。该基准包含817个问题,涵盖38个类别,包括健康、法律、金融和政治。我们精心设计了一些由于错误信念或误解而一些人会答错的问题。为了表现良好,模型必须避免生成从模仿人类文本中学到的错误答案。我们测试了GPT - 3、GPT - Neo/J、GPT - 2以及一个基于T5的模型。最好的模型在58%的问题上是真实的,而人类的表现是94%。模型生成了许多模仿流行误解且有可能欺骗人类的错误答案。最大的模型通常是最不真实的。这与其他自然语言处理任务形成对比,在其他任务中,性能随着模型规模的增大而提高。然而,如果错误答案是从训练分布中学习到的,那么这个结果是意料之中的。我们认为,与使用除模仿网络文本之外的训练目标进行微调相比,仅仅扩大模型规模对于提高真实性的前景不太乐观。
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. We tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model. The best model was truthful on 58% of questions, while human performance was 94%. Models generated many false answers that mimic popular misconceptions and have the potential to deceive humans. The largest models were generally the least truthful. This contrasts with other NLP tasks, where performance improves with model size. However, this result is expected if false answers are learned from the training distribution. We suggest that scaling up models alone is less promising for improving truthfulness than fine-tuning using training objectives other than imitation of text from the web.