Constant interaction-time scatter/gather browsing of very large document collections

Constant interaction-time scatter/gather browsing of very large document collections
复制标题

DOI:
10.1145/160688.160706
复制
发表时间:
1993-07
期刊:
--
影响因子:
--
通讯作者:
D. Cutting;David R Karger;Jan O. Pedersen
D. Cutting;David R Karger;Jan O. Pedersen
中科院分区:
其他
文献类型:
--
作者:
D. Cutting;David R Karger;Jan O. Pedersen

文献摘要

被引文献

相似文献

Scatter/Gather文档浏览方法使用快速文档聚类来生成大型文档集合的类似目录的轮廓。以前的工作[1]开发了线性时间文档聚类算法,以建立这种方法在中等规模集合上的可行性。然而,即使是线性时间算法也太慢,无法支持非常大的集合(如Tipster,DARPA标准文本检索评估集合)的交互式浏览。我们提出了一个方案,支持恒定的交互时间分散/收集任意大的集合后,近线性时间预处理。这涉及到集群层次结构的构建。一个修改的分散/聚集采用此方案,并提出了一个例子,它的使用超过Tipster收集。
The Scatter/Gather document browsing method uses fast document clustering to produce table-of-contents-like outlines of large document collections. Previous work [1] developed linear-time document clustering algorithms to establish the feasibility of this method over moderately large collections. However, even linear-time algorithms are too slow to support interactive browsing of very large collections such as Tipster, the DARPA standard text retrieval evaluation collection. We present a scheme that supports constant interaction-time Scatter/Gather of arbitrarily large collections after near-linear time preprocessing. This involves the construction of a cluster hierarchy. A modification of Scatter/Gather employing this scheme, and an example of its use over the Tipster collection are presented.