Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders

Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders
复制标题

DOI:
10.18653/v1/2021.emnlp-main.109
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Fangyu Liu;Ivan Vulic;A. Korhonen;Nigel Collier
Fangyu Liu;Ivan Vulic;A. Korhonen;Nigel Collier
中科院分区:
其他
文献类型:
--
作者:
Fangyu Liu;Ivan Vulic;A. Korhonen;Nigel Collier

文献摘要

被引文献

相似文献

先前的研究表明,预训练的掩码语言模型(mlm)作为通用的现成词汇和句子编码器并不有效,也就是说,如果没有对NLI、句子相似度或使用注释任务数据的释义任务进行进一步的任务特定微调。在这项工作中,我们证明了即使没有任何额外的数据,仅仅依靠自我监督,也可以将传销转化为有效的词汇和句子编码器。我们提出了一种非常简单,快速,有效的对比学习技术,称为Mirror-BERT,它可以在20-30秒内将传销(例如BERT和RoBERTa)转换为这样的编码器,而无需访问额外的外部知识。Mirror-BERT依赖于相同和稍微修改的字符串对作为正(即同义)微调示例,并旨在在“身份微调”期间最大化它们的相似性。我们报告说,在不同领域和不同语言的词汇级和句子级任务中,使用Mirror-BERT的传销比现成的传销都有巨大的进步。值得注意的是,在句子相似性(STS)和问答蕴涵(QNLI)任务中,我们的自监督镜像bert模型的性能甚至可以与先前依赖于注释任务数据的句子bert模型相匹配。最后,我们深入研究了传销的内部工作原理,并提出了一些证据,说明为什么这种简单的Mirror-BERT微调方法可以产生有效的通用词汇和句子编码器。
Previous work has indicated that pretrained Masked Language Models (MLMs) are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further task-specific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data. In this work, we demonstrate that it is possible to turn MLMs into effective lexical and sentence encoders even without any additional data, relying simply on self-supervision. We propose an extremely simple, fast, and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds with no access to additional external knowledge. Mirror-BERT relies on identical and slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during “identity fine-tuning”. We report huge gains over off-the-shelf MLMs with Mirror-BERT both in lexical-level and in sentence-level tasks, across different domains and different languages. Notably, in sentence similarity (STS) and question-answer entailment (QNLI) tasks, our self-supervised Mirror-BERT model even matches the performance of the Sentence-BERT models from prior work which rely on annotated task data. Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple Mirror-BERT fine-tuning approach can yield effective universal lexical and sentence encoders.