XRANK: ranked keyword search over XML documents

XRANK: ranked keyword search over XML documents
复制标题

DOI:
10.1145/872757.872762
复制
发表时间:
2003-06
期刊:
--
影响因子:
--
通讯作者:
Lin Guo;F. Shao;C. Botev;J. Shanmugasundaram
Lin Guo;F. Shao;C. Botev;J. Shanmugasundaram
中科院分区:
其他
文献类型:
--
作者:
Lin Guo;F. Shao;C. Botev;J. Shanmugasundaram

文献摘要

被引文献

相似文献

我们考虑了通过超链接的XML文档有效产生关键字搜索查询的排名结果的问题。评估关键字搜索查询对层次XML文档的评估,而不是(概念上)flat html文档,引入了许多新的挑战。首先,XML关键字搜索查询并不总是返回整个文档,但是可以返回包含所需关键字的深度嵌套的XML元素。其次,XML的嵌套结构意味着排名的概念不再是文档的粒度,而是在XML元素的粒度上。最后,在层次XML数据模型中,关键字接近度的概念更为复杂。在本文中,我们介绍Xrank系统,该系统旨在处理XML关键字搜索的这些新功能。我们的实验结果表明,与现有方法相比,Xrank具有空间和性能优势。 Xrank的一个有趣功能是,它自然地概括了基于超链接的HTML搜索引擎(例如Google)。因此,Xrank可用于查询HTML和XML文档的混合。
We consider the problem of efficiently producing ranked results for keyword search queries over hyperlinked XML documents. Evaluating keyword search queries over hierarchical XML documents, as opposed to (conceptually) flat HTML documents, introduces many new challenges. First, XML keyword search queries do not always return entire documents, but can return deeply nested XML elements that contain the desired keywords. Second, the nested structure of XML implies that the notion of ranking is no longer at the granularity of a document, but at the granularity of an XML element. Finally, the notion of keyword proximity is more complex in the hierarchical XML data model. In this paper, we present the XRANK system that is designed to handle these novel features of XML keyword search. Our experimental results show that XRANK offers both space and performance benefits when compared with existing approaches. An interesting feature of XRANK is that it naturally generalizes a hyperlink based HTML search engine such as Google. XRANK can thus be used to query a mix of HTML and XML documents.