Crawling the Hidden Web

Crawling the Hidden Web
复制标题

DOI:
--
复制
发表时间:
2001-09
期刊:
--
影响因子:
--
通讯作者:
S. Raghavan;H. Garcia-Molina
S. Raghavan;H. Garcia-Molina
中科院分区:
其他
文献类型:
--
作者:
S. Raghavan;H. Garcia-Molina

文献摘要

被引文献

相似文献

当前的爬虫只从可公开索引的Web检索内容,也就是说,仅通过遵循超文本链接即可访问的Web页面集,忽略需要授权或事先注册的搜索表单和页面。特别是,他们忽略了大量高质量的内容“隐藏”在搜索表单后面,在大型可搜索的电子数据库中。在本文中,我们解决了设计一个能够从这个隐藏的Web中提取内容的爬虫的问题。我们介绍了一个隐藏网络爬虫的通用操作模型,并描述了这个模型是如何在HiWE(隐藏网络暴露器)中实现的,HiWE是斯坦福大学开发的一个原型爬虫。我们介绍了一种新的基于布局的信息提取技术(LITE),并演示了它在从搜索表单和响应页面中自动提取语义信息中的应用。我们还介绍了为测试和验证我们的技术而进行的实验结果。
Current-day crawlers retrieve content only from the publicly indexable Web, i.e., the set of Web pages reachable purely by following hypertext links, ignoring search forms and pages that require authorization or prior registration. In particular, they ignore the tremendous amount of high quality content “hidden” behind search forms, in large searchable electronic databases. In this paper, we address the problem of designing a crawler capable of extracting content from this hidden Web. We introduce a generic operational model of a hidden Web crawler and describe how this model is realized in HiWE (Hidden Web Exposer), a prototype crawler built at Stanford. We introduce a new Layout-based Information Extraction Technique (LITE) and demonstrate its use in automatically extracting semantic information from search forms and response pages. We also present results from experiments conducted to test and validate our techniques.