Structured databases on the web: observations and implications

Structured databases on the web: observations and implications
复制标题

DOI:
10.1145/1031570.1031584
复制
发表时间:
2004-09
期刊:
SIGMOD Rec.
影响因子:
--
通讯作者:
K. Chang;Bin He;Chengkai Li;Mitesh Patel;Zhen Zhang
K. Chang;Bin He;Chengkai Li;Mitesh Patel;Zhen Zhang
中科院分区:
其他
文献类型:
--
作者:
K. Chang;Bin He;Chengkai Li;Mitesh Patel;Zhen Zhang

文献摘要

被引文献

相似文献

互联网已经被在线数据库的普及迅速“深化”。由于潜在的无限信息隐藏在其查询接口后面,这种可搜索数据库的“深层网络”显然是数据访问的重要前沿。本文调查这一相对未开发的前沿,测量相关的特点,探索和整合结构化的Web资源。一方面,我们的“宏观”研究调查了整个深网,在2004年4月,采用随机IP抽样方法,有100万个样本。(How深网是什么?目前的目录服务如何覆盖它?)另一方面,我们的“微观”研究调查源的具体特征超过441个来源在八个代表性的领域,在2002年12月。(How“隐藏的”是深层网络资源吗搜索引擎如何覆盖他们的数据?查询表单的复杂性和表现力如何?)我们报告我们的观察结果,并将结果数据集发布给研究界。我们的结论与几个影响(我们自己的),而一定是主观的,可能有助于形成研究方向和解决方案。
The Web has been rapidly "deepened" by the prevalence of databases online. With the potentially unlimited information hidden behind their query interfaces, this "deep Web" of searchable databses is clearly an important frontier for data access. This paper surveys this relatively unexplored frontier, measuring characteristics pertinent to both exploring and integrating structured Web sources. On one hand, our "macro" study surveys the deep Web at large, in April 2004, adopting the random IP-sampling approach, with one million samples. (How large is the deep Web? How is it covered by current directory services?) On the other hand, our "micro" study surveys source-specific characteristics over 441 sources in eight representative domains, in December 2002. (How "hidden" are deep-Web sources? How do search engines cover their data? How complex and expressive are query forms?) We report our observations and publish the resulting datasets to the research community. We conclude with several implications (of our own) which, while necessarily subjective, might help shape research directions and solutions.