The anatomy of a large-scale hypertextual Web search engine

The anatomy of a large-scale hypertextual Web search engine
复制标题

DOI:
10.1016/s0169-7552(98)00110-x
复制
发表时间:
1998-04-01
期刊:
COMPUTER NETWORKS AND ISDN SYSTEMS
影响因子:
--
通讯作者:
Page, L
Page, L
中科院分区:
其他
文献类型:
--
作者:
Brin, S;Page, L

文献摘要

被引文献

相似文献

在本文中,我们提出了谷歌,一个大规模的搜索引擎,大量使用的超文本结构的原型。Google的设计目的是高效地抓取和索引Web,并产生比现有系统更令人满意的搜索结果。该原型有一个全文和超链接数据库,至少有2400万页,可在http://google.stanford.edu/To上查阅。搜索引擎索引数以千万计的网页,涉及相当数量的不同的条款。他们每天回答数以千万计的问题。尽管大型搜索引擎在网络上的重要性,很少有学术研究已经做了。此外,由于技术的快速发展和Web的激增,今天创建Web搜索引擎与三年前有很大不同。本文深入描述了我们的大型Web搜索引擎--这是我们迄今为止所知道的第一个如此详细的公开描述。除了将传统搜索技术扩展到如此大规模的数据的问题之外,还存在新的技术挑战使用超文本中存在的附加信息来产生更好的搜索结果。本文讨论了如何建立一个实用的大规模系统,可以利用超文本中存在的附加信息的问题。此外,我们看看如何有效地处理不受控制的超文本集合,任何人都可以发布任何他们想要的问题。(C)1998年由Elsevier Science B. V.出版,版权所有。
In this paper, we present Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext. Google is designed to crawl and index the Web efficiently and produce much more satisfying search results than existing systems. The prototype with a full text and hyperlink database of at least 24 million pages is available at http://google.stanford.edu/To engineer a search engine is a challenging task. Search engines index tens to hundreds of millions of Web pages involving a comparable number of distinct terms. They answer tens of millions of queries every day. Despite the importance of large-scale search engines on the Web, very little academic research has been done on them. Furthermore, due to rapid advance in technology and Web proliferation, creating a Web search engine today is very different from three years ago. This paper provides an in-depth description of our large-scale Web search engine - the first such detailed public description we know of to date.Apart from the problems of scaling traditional search techniques to data of this magnitude, there are new technical challenges involved with using the additional information present in hypertext to produce better search results. This paper addresses this question of how to build a practical large-scale system which can exploit the additional information present in hypertext. Also we look at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want. (C) 1998 Published by Elsevier Science B.V. All rights reserved.