Pruning long documents for distributed information retrieval

Pruning long documents for distributed information retrieval
复制标题

修剪长文档以进行分布式信息检索

DOI:
10.1145/584792.584847
复制
发表时间:
2002
期刊:
Lit. Linguistic Comput.
影响因子:
--
通讯作者:
Jamie Callan
Jamie Callan
中科院分区:
--
文献类型:
--
作者:
Jie Lu;Jamie Callan

文献摘要

被引文献

相似文献

基于查询的采样是一种通过向搜索引擎提交查询并观察返回的文档来发现文本数据库内容的方法。在以前的研究中,样本文档被用来建立资源描述,用于自动数据库选择,并建立一个集中的样本数据库,用于查询扩展和结果合并。一个未说明的假设是,相关的存储成本是可以接受的。当样本文档很长时,存储成本可能很大。本文研究了修剪长文档以降低存储成本的方法。实验结果表明,构建资源描述和集中样本数据库的样本文件的修剪内容可以减少存储成本的54-93%,而只造成轻微的损失,在分布式信息检索的准确性。
Query-based sampling is a method of discovering the contents of a text database by submitting queries to a search engine and observing the documents returned. In prior research sampled documents were used to build resource descriptions for automatic database selection, and to build a centralized sample database for query expansion and result merging. An unstated assumption was that the associated storage costs were acceptable.When sampled documents are long, storage costs can be large. This paper investigates methods of pruning long documents to reduce storage costs. The experimental results demonstrate that building resource descriptions and centralized sample databases from the pruned contents of sampled documents can reduce storage costs by 54-93% while causing only minor losses in the accuracy of distributed information retrieval.