Deep Web Structure

Deep Web Structure
复制标题

DOI:
10.1109/mic.2002.1036032
复制
发表时间:
2002-09
期刊:
IEEE Internet Comput.
影响因子:
--
通讯作者:
Munindar P. Singh
Munindar P. Singh
中科院分区:
其他
文献类型:
--
作者:
Munindar P. Singh

文献摘要

被引文献

相似文献

我们目前对Web结构的理解是基于由集中式爬虫和索引器创建的大型图。他们几乎完全从所谓的表层网络获得数据,粗略地说,表层网络由相互链接的HTML页面组成。相比之下,深层网络是可以通过Web访问的信息,但驻留在数据库中;它是动态可用的,以响应查询,而不是提前放置在静态页面上。最近的估计表明,深层网络的数据量是表层网络的数百倍。深层网络让我们有理由重新思考目前广泛的链接分析理论。网络爬虫将不得不产生查询来生成相关页面,而不是查找页面并在页面上找到链接。如果不了解被查询站点的内容,提前创建适当的查询是非常重要的。深网的规模也使得缓存结果比仅仅索引静态页面要困难得多。静态页面将其链接呈现给所有人,而深层Web站点可以决定处理谁的查询以及如何处理。例如,它可以在提供任何真正有价值的信息和链接之前对查询方进行身份验证。它可以了解查询方的上下文,以便给出适当的响应,并且可以进行对话并协商它所揭示的信息。因此,网站可以防止其信息被未知方使用。更重要的是,查询方可以确保信息是为它准备的。
Our current understanding of Web structure is based on large graphs created by centralized crawlers and indexers. They obtain data almost exclusively from the so-called surface Web, which consists, loosely speaking, of interlinked HTML pages. The deep Web, by contrast, is information that is reachable over the Web, but that resides in databases; it is dynamically available in response to queries, not placed on static pages ahead of time. Recent estimates indicate that the deep Web has hundreds of times more data than the surface Web. The deep Web gives us reason to rethink much of the current doctrine of broad-based link analysis. Instead of looking up pages and finding links on them, Web crawlers would have to produce queries to generate relevant pages. Creating appropriate queries ahead of time is nontrivial without understanding the content of the queried sites. The deep Web's scale would also make it much harder to cache results than to merely index static pages. Whereas a static page presents its links for all to see, a deep Web site can decide whose queries to process and how well. It can, for example, authenticate the querying party before giving it any truly valuable information and links. It can build an understanding of the querying party's context in order to give proper responses, and it can engage in dialogues and negotiate for the information it reveals. The Web site can thus prevent its information from being used by unknown parties. What's more, the querying party can ensure that the information is meant for it.