Constant interaction-time scatter/gather browsing of very large document collections
Constant interaction-time scatter/gather browsing of very large document collections
复制标题
DOI:
10.1145/160688.160706
复制
发表时间:
1993-07
期刊:
影响因子:
--
通讯作者:
D. Cutting;David R Karger;Jan O. Pedersen
中科院分区:
文献类型:
--
作者:
D. Cutting;David R Karger;Jan O. Pedersen
The Scatter/Gather document browsing method uses fast document clustering to produce table-of-contents-like outlines of large document collections. Previous work [1] developed linear-time document clustering algorithms to establish the feasibility of this method over moderately large collections. However, even linear-time algorithms are too slow to support interactive browsing of very large collections such as Tipster, the DARPA standard text retrieval evaluation collection. We present a scheme that supports constant interaction-time Scatter/Gather of arbitrarily large collections after near-linear time preprocessing. This involves the construction of a cluster hierarchy. A modification of Scatter/Gather employing this scheme, and an example of its use over the Tipster collection are presented.