THE EFFECTIVENESS OF DOCUMENT NEIGHBORING IN SEARCH ENHANCEMENT

THE EFFECTIVENESS OF DOCUMENT NEIGHBORING IN SEARCH ENHANCEMENT
复制标题

DOI:
10.1016/0306-4573(94)90068-x
复制
发表时间:
1994-03-01
影响因子:
8.6
通讯作者:
COFFEE, L
COFFEE, L
中科院分区:
计算机科学1区
文献类型:
--
作者:
WILBUR, WJ;COFFEE, L

文献摘要

被引文献

相似文献

我们考虑两种可以应用于数据库的查询。第一个是搜索者为表达信息需求而编写的查询。第二个是对与搜索者已经判断相关的文档最相似的文档的请求。我们检查了这两个过程的有效性,并表明在重要情况下后一种查询类型比前一种更有效。这提供了聚类假设的新观点以及文档相邻过程(密切相关文档的预计算)的合理性。如果数据库中的所有文档都有现成的预先计算的最近邻,则可以方便地使用一种新的搜索算法,我们称之为并行邻域搜索。我们表明,这种基于反馈的方法比传统的线性搜索方法在召回率方面提供了显着的改进,甚至在整体性能上优于传统的反馈方法。
We consider two kinds of queries that may be applied to a database. The first is a query written by a searcher to express an information need. The second is a request for documents most similar to a document already judged relevant by the searcher. We examine the effectiveness of these two procedures and show that in important cases the latter query type is more effective than the former. This provides a new view of the cluster hypothesis and a justification for document neighboring procedures (precomputation of closely related documents). If all the documents in a database have readily available precomputed nearest neighbors, a new search algorithm, which we call parallel neighborhood searching, is conveniently used. We show that this feedback-based method provides significant improvement in recall over traditional linear searching methods, and even appears superior to traditional feedback methods in overall performance.