Extracting Patterns and Relations from the World Wide Web

Extracting Patterns and Relations from the World Wide Web
复制标题

DOI:
10.1007/10704656_11
复制
发表时间:
1998-03
期刊:
--
影响因子:
--
通讯作者:
Sergey Brin
Sergey Brin
中科院分区:
其他
文献类型:
--
作者:
Sergey Brin

文献摘要

被引文献

相似文献

万维网是一个巨大的信息资源。与此同时,它是非常分散的。特定类型的数据(例如餐厅列表)可能以许多不同的格式分散在数千个独立的信息源中。在本文中,我们考虑的问题,这样的数据类型自动从所有这些来源提取的关系。我们提出了一种技术,利用模式和关系集之间的二元性,从一个小样本开始增长的目标关系。为了测试我们的技术,我们使用它来提取一个关系(作者,标题)对从万维网。
The World Wide Web is a vast resource for information. At the same time it is extremely distributed. A particular type of data such as restaurant lists may be scattered across thousands of independent information sources in many different formats. In this paper, we consider the problem of extracting a relation for such a data type from all of these sources automatically. We present a technique which exploits the duality between sets of patterns and relations to grow the target relation starting from a small sample. To test our technique we use it to extract a relation of (author,title) pairs from the World Wide Web.