On the temporal dimension of search

On the temporal dimension of search
复制标题

DOI:
10.1145/1013367.1013519
复制
发表时间:
2004-05
期刊:
--
影响因子:
--
通讯作者:
Philip S. Yu;Xin Li;B. Liu
Philip S. Yu;Xin Li;B. Liu
中科院分区:
其他
文献类型:
--
作者:
Philip S. Yu;Xin Li;B. Liu

文献摘要

被引文献

相似文献

网络搜索可能是互联网上最重要的一个应用程序。最著名的搜索技术可能是PageRank和HITS算法。这些算法的动机是观察到从一个页面到另一个页面的超链接是对目标页面的隐含的权威传递。他们利用这一社会现象来识别高质量的页面,例如“权威”页面和“中心”页面。在这篇文章中,我们认为这些算法忽略了Web的一个重要维度,即时间维度。网络不是一个静态的环境。它是不断变化的。过去的高质量页面现在或将来可能不是高质量的页面。这些技术倾向于较老的页面,因为这些页面随着时间的推移积累了许多内链接。新页面可能是高质量的,没有或很少有内嵌链接,因此会被抛在后面。给用户带来新的高质量的页面很重要,因为大多数用户想要最新的信息。研究出版物搜索也存在同样的问题。本文以科研出版物检索为背景,对检索的时间维度进行了研究。我们提出了一些方法来处理这个问题。实验结果表明,这些方法是非常有效的。
Web search is probably the single most important application on the Internet. The most famous search techniques are perhaps the PageRank and HITS algorithms. These algorithms are motivated by the observation that a hyperlink from a page to another is an implicit conveyance of authority to the target page. They exploit this social phenomenon to identify quality pages, e.g., "authority" pages and "hub" pages. In this paper we argue that these algorithms miss an important dimension of the Web, the temporal dimension. The Web is not a static environment. It changes constantly. Quality pages in the past may not be quality pages now or in the future. These techniques favor older pages because these pages have many in-links accumulated over time. New pages, which may be of high quality, have few or no in-links and are left behind. Bringing new and quality pages to users is important because most users want the latest information. Research publication search has exactly the same problem. This paper studies the temporal dimension of search in the context of research publication search. We propose a number of methods deal with the problem. Our experimental results show that these methods are highly effective.