What's really new on the web?: identifying new pages from a series of unstable web snapshots

What's really new on the web?: identifying new pages from a series of unstable web snapshots
复制标题

网络上真正的新鲜事是什么?:从一系列不稳定的网络快照中识别新页面

DOI:
10.1145/1135777.1135815
复制
发表时间:
2006
期刊:
影响因子:
3.7
通讯作者:
M. Kitsuregawa
M. Kitsuregawa
中科院分区:
医学3区
文献类型:
--
作者:
Masashi Toyoda;M. Kitsuregawa

文献摘要

被引文献

相似文献

识别和跟踪网络上的新信息在社会学、市场营销和调查研究中很重要,因为新的趋势可能在新信息中很明显。这种变化可以通过定期抓取Web来观察。然而,在实践中,不可能重复地抓取整个扩展的Web。这意味着页面的新奇仍然未知,即使该页面在之前的快照中不存在。在本文中,我们提出了一个新奇的措施,估计确定性,一个新的抓取页面出现在以前和当前的抓取。使用这种新奇度量,可以从一系列不稳定的快照中提取新页面,以进行进一步的分析和挖掘,从而识别Web上的新趋势。我们评估的精度,召回率和失误率的新奇措施,使用我们的日本网络档案,并将其应用到网络档案搜索引擎。
Identifying and tracking new information on the Web is important in sociology, marketing, and survey research, since new trends might be apparent in the new information. Such changes can be observed by crawling the Web periodically. In practice, however, it is impossible to crawl the entire expanding Web repeatedly. This means that the novelty of a page remains unknown, even if that page did not exist in previous snapshots. In this paper, we propose a novelty measure for estimating the certainty that a newly crawled page appeared between the previous and current crawls. Using this novelty measure, new pages can be extracted from a series of unstable snapshots for further analysis and mining to identify new trends on the Web. We evaluated the precision, recall, and miss rate of the novelty measure using our Japanese web archive, and applied it to a Web archive search engine.