Common Sense or World Knowledge? Investigating Adapter-Based Knowledge Injection into Pretrained Transformers

Common Sense or World Knowledge? Investigating Adapter-Based Knowledge Injection into Pretrained Transformers
复制标题

DOI:
10.18653/v1/2020.deelio-1.5
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Anne Lauscher;Olga Majewska;Leonardo F. R. Ribeiro;Iryna Gurevych;N. Rozanov;Goran Glavavs
Anne Lauscher;Olga Majewska;Leonardo F. R. Ribeiro;Iryna Gurevych;N. Rozanov;Goran Glavavs
中科院分区:
其他
文献类型:
--
作者:
Anne Lauscher;Olga Majewska;Leonardo F. R. Ribeiro;Iryna Gurevych;N. Rozanov;Goran Glavavs

文献摘要

相似文献

在BERT或GPT-2等神经语言模型(LM)在各种语言理解任务上取得重大成功之后,最近的工作集中在将来自外部资源的(结构化)知识注入这些模型中。一方面,联合预先训练(即,从头开始训练,将基于外部知识的目标添加到主要LM目标)可能在计算上昂贵得令人望而却步,另一方面,对外部知识的事后微调可能导致对分布式知识的灾难性遗忘。在这项工作中,我们调查模型补充BERT的分布知识与概念知识从ConceptNet及其相应的开放式思维常识(OMCS)语料库,分别使用适配器训练。虽然GLUE基准测试的总体结果并不确定,但更深入的分析表明,我们基于适配器的模型在推理任务上的表现大大优于BERT(高达15-20个性能点),这些推理任务需要ConceptNet和OMCS中明确存在的概念知识类型。我们还在https://github.com/wluper/retrograph下开源了我们所有的实验和相关代码。
Following the major success of neural language models (LMs) such as BERT or GPT-2 on a variety of language understanding tasks, recent work focused on injecting (structured) knowledge from external resources into these models. While on the one hand, joint pre-training (i.e., training from scratch, adding objectives based on external knowledge to the primary LM objective) may be prohibitively computationally expensive, post-hoc fine-tuning on external knowledge, on the other hand, may lead to the catastrophic forgetting of distributional knowledge. In this work, we investigate models for complementing the distributional knowledge of BERT with conceptual knowledge from ConceptNet and its corresponding Open Mind Common Sense (OMCS) corpus, respectively, using adapter training. While overall results on the GLUE benchmark paint an inconclusive picture, a deeper analysis reveals that our adapter-based models substantially outperform BERT (up to 15-20 performance points) on inference tasks that require the type of conceptual knowledge explicitly present in ConceptNet and OMCS. We also open source all our experiments and relevant code under: https://github.com/wluper/retrograph.