A Multimodal Translation-Based Approach for Knowledge Graph Representation Learning

A Multimodal Translation-Based Approach for Knowledge Graph Representation Learning
复制标题

DOI:
10.18653/v1/s18-2027
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
Hatem Mousselly Sergieh;Teresa Botschen;Iryna Gurevych;S. Roth
Hatem Mousselly Sergieh;Teresa Botschen;Iryna Gurevych;S. Roth
中科院分区:
其他
文献类型:
--
作者:
Hatem Mousselly Sergieh;Teresa Botschen;Iryna Gurevych;S. Roth

文献摘要

相似文献

目前的知识图表示学习方法只关注知识图的结构,而不利用任何类型的外部信息,例如与知识图实体相对应的视觉和语言信息。在本文中,我们提出了一种基于多模态分析的方法,该方法将KG三元组的能量定义为利用多模态(视觉和语言)和结构KG表示的子能量函数的总和。接下来,使用简单的神经网络架构来最小化基于排名的损失。此外,我们引入了一个新的大规模数据集用于多模态KG表示学习。我们比较了我们的方法与其他基线在两个标准任务上的性能,即知识图完成和三重分类,使用我们的以及WN 9-IMG数据集。结果表明,我们的方法在任务和数据集上都优于所有基线。
Current methods for knowledge graph (KG) representation learning focus solely on the structure of the KG and do not exploit any kind of external information, such as visual and linguistic information corresponding to the KG entities. In this paper, we propose a multimodal translation-based approach that defines the energy of a KG triple as the sum of sub-energy functions that leverage both multimodal (visual and linguistic) and structural KG representations. Next, a ranking-based loss is minimized using a simple neural network architecture. Moreover, we introduce a new large-scale dataset for multimodal KG representation learning. We compared the performance of our approach to other baselines on two standard tasks, namely knowledge graph completion and triple classification, using our as well as the WN9-IMG dataset. The results demonstrate that our approach outperforms all baselines on both tasks and datasets.