A Ranking Scheme for XML Information Retrieval Based on Benefit and Reading Effort

A Ranking Scheme for XML Information Retrieval Based on Benefit and Reading Effort
复制标题

DOI:
10.1007/978-3-540-77094-7_32
复制
发表时间:
2007-12
期刊:
--
影响因子:
--
通讯作者:
Toshiyuki Shimizu;Masatoshi Yoshikawa
Toshiyuki Shimizu;Masatoshi Yoshikawa
中科院分区:
其他
文献类型:
--
作者:
Toshiyuki Shimizu;Masatoshi Yoshikawa

文献摘要

相似文献

XML信息检索(XML-IR)系统搜索给定查询的XML文档中的相关文档片段。在top-kearch中,用户通过一个整数来控制输出的大小。然而,在XML-IR中,每个输出元素的大小差别很大。因此,通过简单地给出一个整数k,顶层k元素的总输出大小是无法控制的。此外,搜索结果可能包含嵌套元素。如果系统仅根据结果元素的相关性对其进行排序,则由于嵌套,我们可能会多次浏览相同的内容。为了解决这些问题,我们提出了一种新的排序方法,通过引入效益和阅读努力的概念,使我们能够有效地浏览XML-IR系统的搜索结果。提出了一种基于收益和阅读努力的评价指标,并通过实验与现有的XML-IR评价指标进行了比较。
XML information retrieval (XML-IR) systems search for relevant document fragments in XML documents for given queries. In top-ksearch, users control the size of output by an integerk. In XML-IR, however, each output element varies widely in size. Consequently, total output size of top-kelements is uncontrollable by simply giving an integerk. In addition, search results may have nesting elements. If a system orders result elements simply by their relevance, we may browse the same content more than once due to the nestings. To handle these problems, we propose a new ranking method that enables us to browse search results of XML-IR systems efficiently by introducing the concepts ofbenefitandreading effort. We also propose an evaluation metrics based onbenefitandreading effort, and compared the metrics with existing XML-IR metrics by experiments.