Human Performance on Clustering Web Pages

Human Performance on Clustering Web Pages
复制标题

聚类网页上的人类表现

DOI:
10.7282/t3pv6pz1
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
H. Hirsh
H. Hirsh
中科院分区:
--
文献类型:
--
作者:
Sofus A. Macskassy;Arunava Banerjee;Brian D. Davison;H. Hirsh

文献摘要

被引文献

相似文献

随着万维网上信息的增加,在不使用多个查询或使用特定主题的搜索引擎的情况下快速找到所需信息变得困难。帮助搜索的一种方法是将以某种方式显示为相关的HTML页面分组在一起。为了更好地理解这项任务,我们对网页的人类聚类进行了初步研究,希望它能为自动化这项任务的难度提供一些见解。我们的研究结果表明,受试者并没有集群相同,事实上,平均而言,任何两个科目在他们的网页集群的相似性很小。我们还发现,主题通常会创建相当小的集群,而那些只访问URL的主题比那些访问每个网页全文的主题创建的集群更少。一般来说,当给出全文时,任何给定主题的集群之间的文件重叠增加,集群文件的百分比也增加。当分析单个主题时,我们发现每个主题在查询中有不同的行为,无论是在重叠,集群大小还是集群数量方面。这些结果为任何寻求一种明确正确的网页聚类方法提供了一个清醒的说明。
With the increase in information on the World Wide Web it has become difficult to find desired information quickly without using multiple queries or using a topic-specific search engine. One way to help in the search is by grouping HTML pages together that appear in some way to be related. In order to better understand this task, we performed an initial study of human clustering of web pages, in the hope that it would provide some insight into the difficulty of automating this task. Our results show that subjects did not cluster identically; in fact, on average, any two subjects had little similarity in their web-page clusters. We also found that subjects generally created rather small clusters, and those with access only to URLs created fewer clusters than those with access to the full text of each web page. Generally the overlap of documents between clusters for any given subject increased when given the full text, as did the percentage of documents clustered. When analyzing individual subjects, we found that each had different behavior across queries, both in terms of overlap, size of clusters, and number of clusters. These results provide a sobering note on any quest for a single clearly correct clustering method for web pages.1