Pruning long documents for distributed information retrieval
Pruning long documents for distributed information retrieval
复制标题
修剪长文档以进行分布式信息检索
DOI:
10.1145/584792.584847
复制
发表时间:
2002
期刊:
影响因子:
--
通讯作者:
Jamie Callan
中科院分区:
文献类型:
--
作者:
Jie Lu;Jamie Callan
Query-based sampling is a method of discovering the contents of a text database by submitting queries to a search engine and observing the documents returned. In prior research sampled documents were used to build resource descriptions for automatic database selection, and to build a centralized sample database for query expansion and result merging. An unstated assumption was that the associated storage costs were acceptable.When sampled documents are long, storage costs can be large. This paper investigates methods of pruning long documents to reduce storage costs. The experimental results demonstrate that building resource descriptions and centralized sample databases from the pruned contents of sampled documents can reduce storage costs by 54-93% while causing only minor losses in the accuracy of distributed information retrieval.