Reducing Non-Normative Text Generation from Language Models

Reducing Non-Normative Text Generation from Language Models
复制标题

DOI:
10.18653/v1/2020.inlg-1.43
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Xiangyu Peng;Siyan Li;Spencer Frazier;Mark O. Riedl
Xiangyu Peng;Siyan Li;Spencer Frazier;Mark O. Riedl
中科院分区:
其他
文献类型:
--
作者:
Xiangyu Peng;Siyan Li;Spencer Frazier;Mark O. Riedl

文献摘要

被引文献

相似文献

GPT-2 等基于 Transformer 的大规模语言模型是在从互联网上抓取的各种语料库上进行预训练的。因此,它们很容易生成非规范文本(即违反社会规范)。我们引入了一种微调 GPT-2 的技术,使用策略梯度强化学习技术和规范文本分类器来产生奖励和惩罚值。我们使用自动化和人类参与实验在五个数据集上评估我们的技术。与人类对规范和非规范生成文本的黄金标准判断相比,规范文本分类器的准确度为 81-90%。我们的规范微调技术能够将非规范文本减少 27-61%,具体取决于数据集。
Large-scale, transformer-based language models such as GPT-2 are pretrained on diverse corpora scraped from the internet. Consequently, they are prone to generating non-normative text (i.e. in violation of social norms). We introduce a technique for fine-tuning GPT-2, using a policy gradient reinforcement learning technique and a normative text classifier to produce reward and punishment values. We evaluate our technique on five data sets using automated and human participant experiments. The normative text classifier is 81-90% accurate when compared to gold-standard human judgements of normative and non-normative generated text. Our normative fine-tuning technique is able to reduce non-normative text by 27-61%, depending on the data set.