RelBERT: Embedding Relations with Language Models

RelBERT: Embedding Relations with Language Models
复制标题

DOI:
10.48550/arxiv.2310.00299
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Asahi Ushio;José Camacho-Collados;Steven Schockaert
Asahi Ushio;José Camacho-Collados;Steven Schockaert
中科院分区:
其他
文献类型:
--
作者:
Asahi Ushio;José Camacho-Collados;Steven Schockaert

文献摘要

相似文献

许多应用程序需要访问有关不同概念和实体如何相关的背景知识。虽然知识图(KG)和大型语言模型(LLM)可以在一定程度上解决这一需求,但KG不可避免地是不完整的,它们的关系模式往往过于粗粒度,而LLM效率低下,难以控制。作为替代方案,我们建议从相对较小的语言模型中提取关系嵌入。特别是,我们表明,掩蔽的语言模型,如RoberTa可以直接微调为这个目的,只使用少量的训练数据。由此产生的模型,我们称之为RelBERT,以令人惊讶的细粒度方式捕获关系相似性,使我们能够在类比基准中设置一个新的最先进的状态。至关重要的是,RelBERT能够建模的关系远远超出了模型在训练过程中所看到的。例如,我们在命名实体之间的关系上获得了强有力的结果,该模型仅在概念之间的词汇关系上进行训练,并且我们观察到RelBERT可以识别形态类比,尽管没有在这些示例上进行训练。总体而言,我们发现RelBERT显著优于基于提示语言模型的策略,这些模型大了几个数量级,包括最近基于GPT的模型和开源模型。
Many applications need access to background knowledge about how different concepts and entities are related. Although Knowledge Graphs (KG) and Large Language Models (LLM) can address this need to some extent, KGs are inevitably incomplete and their relational schema is often too coarse-grained, while LLMs are inefficient and difficult to control. As an alternative, we propose to extract relation embeddings from relatively small language models. In particular, we show that masked language models such as RoBERTa can be straightforwardly fine-tuned for this purpose, using only a small amount of training data. The resulting model, which we call RelBERT, captures relational similarity in a surprisingly fine-grained way, allowing us to set a new state-of-the-art in analogy benchmarks. Crucially, RelBERT is capable of modelling relations that go well beyond what the model has seen during training. For instance, we obtained strong results on relations between named entities with a model that was only trained on lexical relations between concepts, and we observed that RelBERT can recognise morphological analogies despite not being trained on such examples. Overall, we find that RelBERT significantly outperforms strategies based on prompting language models that are several orders of magnitude larger, including recent GPT-based models and open source models.