BUS: an effective indexing and retrieval scheme in structured documents
BUS: an effective indexing and retrieval scheme in structured documents
复制标题
BUS:结构化文档中有效的索引和检索方案
DOI:
10.1145/276675.276702
复制
发表时间:
1998
期刊:
影响因子:
--
通讯作者:
Honglan Jin
中科院分区:
文献类型:
--
作者:
Dongwook Shin;H. Jang;Honglan Jin
In recent digital library systems or World Wide Web environment, many documents are beginning to be provided in the structured format, tagged in mark up languages like SGML or XML. Hence, indexing and query evaluation of structured documents have been drawing attention since they enable to access and retrieve a certain part of documents easily. However, conventional information retrieval techniques do not scale up well in structured documents. This paper suggests an efficient indexing and query evaluation scheme for structured documents (named BUS) that minimizes the indexing overhead and guarantees fast query processing at any level in the document structure. The basic idea is that indexing is performed at the lowest level of the given structure and query evaluation computes the similarity at higher level by accumulating the term frequencies at the lowest level in the bottom up way. The accumulators summing up the similarity play the role of accumulating all the term frequencies of the related part at a certain level. This paper also addresses the implementation of BUS and proves that BUS works correctly. In addition, along with several experiments, it shows that BUS facilitates efficient indexing in terms of space and time and guarantees the reasonable retrieval time in response to user queries.