BUS: an effective indexing and retrieval scheme in structured documents

BUS: an effective indexing and retrieval scheme in structured documents
复制标题

BUS:结构化文档中有效的索引和检索方案

DOI:
10.1145/276675.276702
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
Honglan Jin
Honglan Jin
中科院分区:
--
文献类型:
--
作者:
Dongwook Shin;H. Jang;Honglan Jin

文献摘要

被引文献

相似文献

在最近的数字图书馆系统或万维网环境中,许多文档开始以结构化格式提供,以标记语言如SGML或XML标记。因此,结构化文档的索引和查询评估一直受到关注,因为它们可以方便地访问和检索文档的某一部分。然而,传统的信息检索技术在结构化文档中不能很好地扩展。本文提出了一种有效的索引和查询评估计划的结构化文档(命名总线),最大限度地减少索引开销,并保证快速查询处理在任何级别的文档结构。其基本思想是在给定结构的最低层进行索引,查询评估通过自底向上的方式累积最低层的术语频率来计算更高层的相似度。相似度累加器的作用是将所有相关部分的词频累加到一定水平。本文还讨论了总线的实现,并证明了总线的正确工作。此外,沿着的几个实验表明,总线有利于有效的索引在空间和时间方面,并保证了合理的检索时间,以响应用户的查询。
In recent digital library systems or World Wide Web environment, many documents are beginning to be provided in the structured format, tagged in mark up languages like SGML or XML. Hence, indexing and query evaluation of structured documents have been drawing attention since they enable to access and retrieve a certain part of documents easily. However, conventional information retrieval techniques do not scale up well in structured documents. This paper suggests an efficient indexing and query evaluation scheme for structured documents (named BUS) that minimizes the indexing overhead and guarantees fast query processing at any level in the document structure. The basic idea is that indexing is performed at the lowest level of the given structure and query evaluation computes the similarity at higher level by accumulating the term frequencies at the lowest level in the bottom up way. The accumulators summing up the similarity play the role of accumulating all the term frequencies of the related part at a certain level. This paper also addresses the implementation of BUS and proves that BUS works correctly. In addition, along with several experiments, it shows that BUS facilitates efficient indexing in terms of space and time and guarantees the reasonable retrieval time in response to user queries.