Embedding Multimodal Relational Data for Knowledge Base Completion

Embedding Multimodal Relational Data for Knowledge Base Completion
复制标题

DOI:
10.18653/v1/d18-1359
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Pouya Pezeshkpour;Liyan Chen;Sameer Singh
Pouya Pezeshkpour;Liyan Chen;Sameer Singh
中科院分区:
其他
文献类型:
--
作者:
Pouya Pezeshkpour;Liyan Chen;Sameer Singh

文献摘要

被引文献

相似文献

在嵌入空间中表示实体和关系是关系数据上机器学习的一种经过充分研究的方法。然而,现有的方法,主要集中在一个有限的实体集之间的简单链接结构,忽略了各种数据类型,经常使用的知识库,如文本,图像和数值。在本文中,我们提出了多模态知识库嵌入(MKBE),使用不同的神经编码器的各种观察到的数据,并联合收割机将它们与现有的关系模型学习嵌入的实体和多模态数据。此外,使用这些学习的嵌入和不同的神经解码器,我们引入了一种新的多模态插补模型,从知识库中的信息中生成缺失的多模态值,如文本和图像。我们丰富现有的关系数据集,创建两个新的基准,包含额外的信息,如文本描述和原始实体的图像。我们证明了我们的模型有效地利用了这些额外的信息来提供更准确的链接预测,实现了最先进的结果,与现有方法相比有5-7%的差距。此外,我们通过用户研究来评估我们生成的多模态值的质量。
Representing entities and relations in an embedding space is a well-studied approach for machine learning on relational data. Existing approaches, however, primarily focus on simple link structure between a finite set of entities, ignoring the variety of data types that are often used in knowledge bases, such as text, images, and numerical values. In this paper, we propose multimodal knowledge base embeddings (MKBE) that use different neural encoders for this variety of observed data, and combine them with existing relational models to learn embeddings of the entities and multimodal data. Further, using these learned embedings and different neural decoders, we introduce a novel multimodal imputation model to generate missing multimodal values, like text and images, from information in the knowledge base. We enrich existing relational datasets to create two novel benchmarks that contain additional information such as textual descriptions and images of the original entities. We demonstrate that our models utilize this additional information effectively to provide more accurate link prediction, achieving state-of-the-art results with a considerable gap of 5-7% over existing methods. Further, we evaluate the quality of our generated multimodal values via a user study.