Diverse and proportional size-l object summaries using pairwise relevance

Diverse and proportional size-l object summaries using pairwise relevance
复制标题

DOI:
10.1007/s00778-016-0433-6
复制
发表时间:
2016-07
期刊:
The VLDB Journal
影响因子:
--
通讯作者:
G. Fakas;Zhi Cai;N. Mamoulis
G. Fakas;Zhi Cai;N. Mamoulis
中科院分区:
其他
文献类型:
--
作者:
G. Fakas;Zhi Cai;N. Mamoulis

文献摘要

被引文献

相似文献

图的丰富性和普遍性(例如,在线社交网络,如Google和Facebook;书目图,如DBLP)需要有效和高效的搜索。给定一组可以识别数据主题(DS)的关键字,最近提出的关键字搜索范式产生一组对象摘要(OS)作为结果。OS是以DS节点为根的树结构(即,包含关键字的节点)与周围的节点,这些节点概括了图上保存的关于DS的所有数据。也已经研究了被表示为大小10 S的0 S片段。size-lOS是包含l节点的部分OS,使得它们的重要性分数的总和导致最大可能的总分数。然而,使总重要性得分最大化的节点集合可能导致无信息大小的10 S,因为非常重要的节点可能在其中重复,从而支配其他代表性信息。鉴于这一限制,在本文中,我们研究了两种新型操作系统片段的有效和高效的生成,即,不同的和成比例的大小10 S,表示为DS大小和PSize-10 S。也就是说,除了每个节点的重要性之外,我们还考虑其与OS和片段中的其他节点的成对相关性(相似性)。我们对两个真实的图(DBLP和Google)进行了广泛的评估。我们通过收集用户反馈来验证有效性,例如,通过询问DBLP作者(即,我们的目标是评估我们的结果。此外,我们还验证了算法的效率,并评估了它们产生的片段的质量。
The abundance and ubiquity of graphs (e.g., online social networks such as Googleand Facebook; bibliographic graphs such as DBLP) necessitates the effective and efficient search over them. Given a set of keywords that can identify a data subject (DS), a recently proposed keyword search paradigm produces a set of object summaries (OSs) as results. An OS is a tree structure rooted at the DS node (i.e., a node containing the keywords) with surrounding nodes that summarize all data held on the graph about the DS. OS snippets, denoted as size-lOSs, have also been investigated. A size-lOS is a partial OS containinglnodes such that the summation of their importance scores results in the maximum possible total score. However, the set of nodes that maximize the total importance score may result in an uninformative size-lOSs, as very important nodes may be repeated in it, dominating other representative information. In view of this limitation, in this paper, we investigate the effective and efficient generation of two novel types of OS snippets, i.e.,diverseandproportionalsize-lOSs, denoted asDSize-landPSize-lOSs. Namely, besides the importance of each node, we also consider its pairwise relevance (similarity) to the other nodes in the OS and the snippet. We conduct an extensive evaluation on two real graphs (DBLP and Google). We verify effectiveness by collecting user feedback, e.g., by asking DBLP authors (i.e., the DSs themselves) to evaluate our results. In addition, we verify the efficiency of our algorithms and evaluate the quality of the snippets that they produce.