Entity-Based Language Model Smoothing Approach for Smart Search

Entity-Based Language Model Smoothing Approach for Smart Search
复制标题

DOI:
10.1109/access.2017.2788417
复制
发表时间:
2018
期刊:
影响因子:
3.9
通讯作者:
Feng Zhao;Zeliang Tian;Hai Jin
Feng Zhao;Zeliang Tian;Hai Jin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Feng Zhao;Zeliang Tian;Hai Jin

文献摘要

相似文献

智能搜索在各行各业都发挥着重要作用,例如,根据业务需求,从海量资源中精准搜索所需知识,是提升产业智能化的重要途径。语言模型的平滑对于获得高质量的搜索结果至关重要,因为它有助于减少由数据稀疏引起的不匹配和过拟合问题。传统的平滑方法只关注全局语料库,对文档信息进行局部聚类,缺乏语义分析,导致查询语句和文档之间缺乏语义相关性。在本文中,我们提出了一种基于实体的语言模型平滑智能搜索方法,使用语义相关性,并以实体为桥梁,建立实体语义语言模型,使用知识库。在这种方法中,文档中的实体被链接到外部知识库,例如维基百科。然后,采用软融合和硬融合的方法生成实体语义语言模型。同时,本文还提出了一种两级合并策略,结合Dir-smoothing和JM-smoothing两种方法,根据词与文档的语义相关性对语言模型进行平滑处理。实验结果表明,平滑后的语言模型更接近文档语义主题下的词概率分布,更准确地估计了查询与文档之间的相关性。
Smart search plays an important role in all walks of life, for example, according to business needs, accurate search of required knowledge from massive resources is an important way to enhance industrial intelligence. Smoothing of the language model is essential for obtaining high-quality search results because it helps to reduce mismatching and overfitting problems caused by data sparseness. Traditional smoothing methods lexically focus on the global corpus and locally cluster documents information without semantic analysis, which leads to deficiency of the semantic correlations between query statements and documents. In this paper, we propose an entity-based language model smoothing approach for smart search that uses semantic correlation and takes entities as bridges to build the entity semantic language model using a knowledge base. In this approach, entities in the documents are linked to an external knowledge base, such as Wikipedia. Then, the entity semantic language model is generated by using soft-fused and hard-fused methods. A two-level merging strategy is also presented to smooth the language model according to whether a given word is semantically relevant to the document or not, which integrates the Dir-smoothing and JM-smoothing methods. Experimental results show that the smoothed language model more closely approximates the word probability distribution under the document semantic theme and more accurately estimates the relevance between query and document.