Feedback-driven result ranking and query refinement for exploring semi-structured data collections

Feedback-driven result ranking and query refinement for exploring semi-structured data collections
复制标题

用于探索半结构化数据集合的反馈驱动的结果排名和查询细化

DOI:
10.1145/1739041.1739046
复制
发表时间:
2010
期刊:
Discret. Math.
影响因子:
--
通讯作者:
M. Sapino
M. Sapino
中科院分区:
--
文献类型:
--
作者:
H. Cao;Yan Qi;K. Candan;M. Sapino

文献摘要

被引文献

相似文献

反馈过程在以文档为中心的应用中得到了广泛的应用,如文本检索和多媒体检索。最近,人们也在努力将反馈应用于半结构化的XML文档集合。在本文中,我们注意到反馈也可以成为探索(通过结果排序和查询细化)大型半结构化数据集合的有效工具。特别是,在大规模数据共享和管理环境中,用户可能不知道数据的结构,查询最初可能过于模糊。给定一个路径查询和系统对该数据查询识别的一组结果,我们考虑两种类型的反馈:软反馈捕获用户对某些特性的偏好。另一方面,硬反馈表达了用户关于某些功能是否应该进一步执行或相反应该避免的断言。软反馈和硬反馈都可以是“积极的”或“消极的”。对于软反馈,我们开发了一个概率特征显著性度量,并描述了如何在路径特征之间存在依赖关系的情况下使用它对结果进行排序。为了有效地处理硬反馈(即,足够快地进行交互式探索),我们提出了基于有限自动机的查询优化解决方案。特别地,我们提出了一种新的用于管理硬反馈的LazyDFA+算法。我们还描述了利用反馈过程固有迭代性质的优化。我们将这些技术结合在AXP中,这是一个自适应和探索性路径检索系统。实验结果表明了所提方法的有效性。
Feedback process has been used extensively in document-centric applications, such as text retrieval and multimedia retrieval. Recently, there have been efforts to apply feedback to semi-structured XML document collections as well. In this paper, we note that feedback can also be an effective tool for exploring (through result ranking and query refinement) large semi-structured data collections. In particular, in large scale data sharing and curation environments, where the user may not know the structure of the data, queries may initially be overly vague. Given a path query and a set of results identified by the system to this query over the data, we consider two types of feedback: Soft feedback captures the user's preference for some features over the others. Hard feedback, on the other hand, expresses users' assertions regarding whether certain features should be further enforced or, in contrast, are to be avoided. Both soft and hard feedback can be "positive" or "negative". For soft feedback, we develop a probabilistic feature significance measure and describe how to use this for ranking results in the presence of dependencies between the path features. To deal with the hard feedback efficiently (i.e., fast enough for interactive exploration), we present finite automata based query refinement solutions. In particular, we present a novel LazyDFA+ algorithm for managing hard feedback. We also describe optimizations that leverage the inherently iterative nature of the feedback process. We bring together these techniques in AXP, a system for adaptive and exploratory path retrieval. The experimental results show the effectiveness of the proposed techniques.