gMatch: Knowledge base question answering via semantic matching

gMatch: Knowledge base question answering via semantic matching
复制标题

gMatch:通过语义匹配进行知识库问答

DOI:
10.1016/j.knosys.2021.107270
复制
发表时间:
2021-07-06
影响因子:
8.8
通讯作者:
Wang, Junhu
Wang, Junhu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jiao, Jie;Wang, Shujun;Wang, Junhu

文献摘要

被引文献

相似文献

有效性是知识库问答(KBQA)的关键,以确定查询是否可以返回正确的答案。现有的KBQA工作主要集中在将输入问题转换为相应的逻辑格式,如SPARQL查询。然而,由于这些工作在很大程度上与知识库解耦,因此转换后的查询可能是无效的。在本文中,我们提出了一种新的基于语义匹配的方法,通过提取知识库的子图来建模输入问题的查询意图。将SPARQL查询的生成归结为知识库中的语义匹配,解决了查询的无效性。首先,提出了一种语义查询图来建模输入问题的可靠查询意图。通过对知识库中的语义查询图进行匹配,提取SPARQL查询图。其次,开发了一种基于嵌入的方法来在公共空间中表示不同形式的问题和查询。使用公共表示可以很容易地检测问题和转换后的查询之间的语义丢失。最后,提出了一种数据驱动的语义补全技术,通过在知识库中扩展不完整的SPARQL查询来减少语义损失。在基准数据集上的实验表明,该方法在效率和有效性方面明显优于现有方法。(C)2021年由Elsevier B.V.出版
Effectiveness is essential for knowledge base question answering (KBQA) to determine whether the query can return the correct answers. Existing works for KBQA mainly focus on converting input questions into corresponding logic formats, such as SPARQL queries. However, since these works are largely decoupled from the knowledge base, the converted query may be ineffective. In this paper, we propose a novel semantic matching-based approach to model the query intention of the input question by extracting the subgraph of the knowledge base. The generation of the SPARQL query is reduced to semantic matching in the knowledge base to solve the ineffectiveness of the query. Firstly, a semantic query graph is proposed to model the reliable query intention of the input question. The SPARQL query graph could be extracted by matching the semantic query graph in the knowledge base. Secondly, an embedding-based method is developed to represent different forms of questions and queries in a common space. It is easy to detect semantic loss between the question and the converted query with the common representation. Finally, a data-driven semantic completion technique is presented to reduce the semantic loss by expanding the incomplete SPARQL query in the knowledge base. The experiments evaluated on benchmark datasets show that the proposed approach significantly outperforms state-of-the-art methods in efficiency and effectiveness. (C) 2021 Published by Elsevier B.V.