Masked Language Model Scoring

Masked Language Model Scoring
复制标题

DOI:
10.18653/v1/2020.acl-main.240
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Julian Salazar;Davis Liang;Toan Q. Nguyen;K. Kirchhoff
Julian Salazar;Davis Liang;Toan Q. Nguyen;K. Kirchhoff
中科院分区:
其他
文献类型:
--
作者:
Julian Salazar;Davis Liang;Toan Q. Nguyen;K. Kirchhoff

文献摘要

被引文献

相似文献

预训练的掩蔽语言模型(MLM)需要对大多数NLP任务进行精细调整。取而代之的是,我们通过伪对数似然分数(PLL)来评估MLM,这些分数是通过逐个掩蔽令牌来计算的。我们发现,PLL在各种任务中的表现优于GPT-2等自回归语言模型的分数。通过对ASR和NMT假设的研究,Roberta将端到端LibriSpeech模型的WER相对减少了30%,并在低资源翻译对的最新基线上将BLEU的WER加起来达到+1.7%,并从领域适应中进一步获益。我们将这一成功归功于PLL对语言可接受性的无监督表达,没有从左到右的偏见,大大提高了GPT-2的分数(岛屿效应+10分,软式飞艇NPI许可)。人们可以微调MLMS以在没有掩蔽的情况下给出分数,从而能够在单个推理过程中进行计算。总而言之,PLL及其相关的伪困惑(PPPL)使越来越多的预先训练的MLM能够即插即用地使用;例如,我们使用单一的跨语言模型来重新计算多语言的翻译。我们在https://github.com/awslabs/mlm-scoring.发布了我们的语言模型评分库
Pretrained masked language models (MLMs) require finetuning for most NLP tasks. Instead, we evaluate MLMs out of the box via their pseudo-log-likelihood scores (PLLs), which are computed by masking tokens one by one. We show that PLLs outperform scores from autoregressive language models like GPT-2 in a variety of tasks. By rescoring ASR and NMT hypotheses, RoBERTa reduces an end-to-end LibriSpeech model’s WER by 30% relative and adds up to +1.7 BLEU on state-of-the-art baselines for low-resource translation pairs, with further gains from domain adaptation. We attribute this success to PLL’s unsupervised expression of linguistic acceptability without a left-to-right bias, greatly improving on scores from GPT-2 (+10 points on island effects, NPI licensing in BLiMP). One can finetune MLMs to give scores without masking, enabling computation in a single inference pass. In all, PLLs and their associated pseudo-perplexities (PPPLs) enable plug-and-play use of the growing number of pretrained MLMs; e.g., we use a single cross-lingual model to rescore translations in multiple languages. We release our library for language model scoring at https://github.com/awslabs/mlm-scoring.