Result Set Diversification in Digital Libraries Through the Use of Paper's Claims

Result Set Diversification in Digital Libraries Through the Use of Paper's Claims
复制标题

DOI:
10.1007/978-3-319-70232-2_19
复制
发表时间:
2017-11
影响因子:
--
通讯作者:
J. M. Pinto;Wolf-Tilo Balke
J. M. Pinto;Wolf-Tilo Balke
中科院分区:
--
文献类型:
--
作者:
J. M. Pinto;Wolf-Tilo Balke

文献摘要

相似文献

从查询中了解两个实体之间可能的关联是一个难题。例如,即使在精心策划的数字图书馆中查询“咖啡”和“癌症”,对于难以弄清楚查询意图的检索系统来说也是一个挑战。也许用户想要对其已知的内容达成共识?但存在多少种不同的协会呢?如何找到全部?在这里,我们介绍一种使从此类查询检索到的结果多样化的方法,旨在对结果列表进行重新排名。我们的重新排名模型专门针对科学论文的一个基本方面:声明。声明是科学家用来报告发现的句子。特别是,我们研究表达医学领域实体之间关联的主张。更具体地说,我们关注涉及两个实体的查询,其中一个实体对某种疾病有一定影响。因此,我们使用通过查询 PubMed 获得的语料库来根据经验评估我们提出的解决方案。此外,我们将声明的想法推广为考虑查询结果集多样化的明确关键方面。我们展示了我们的方法在简化发现实体之间代表性关联的过程方面的潜力。我们的方法依赖于使用词向量的神经嵌入来表示声明,并实现一种算法来对查询结果集进行重新排序。我们凭经验展示了我们方法的潜力。
Understanding the possible associations between two entities from a query is a hard problem. For instance, querying “coffee” and “cancer” even in a curated Digital Library is a challenge to the retrieval system that struggles to figure out the intention of the query. Maybe the user wants a consensus of what it is known? But how many different associations exist? How to find them all? Herein we introduce an approach to diversify the results retrieved from such queries aiming at re-ranking the result list. Our re-ranking models specifically one fundamental aspect of scientific papers: claims. Claims are the sentences that scientists use to report findings. In particular, we study claims that express associations between entities in the medical domain. More specifically, we focus on queries that involve two entities in which one of the entities has some effect on a disease. Thus, we work on a corpus obtained by querying PubMed to empirically assess our proposed solution. Moreover, we promote the idea of claims as an explicit key aspect to consider diversification in the result set of a query. We show the potential of our approach to ease the process of discovering representative associations between entities. Our approach relies on a representation of claims using neural embedding of word vectors and implements an algorithm to perform the re-ranking of the result set of a query. We empirically show the potential of our approach.