Extraction and integration of web data by end-users
Extraction and integration of web data by end-users
复制标题
最终用户提取和整合网络数据
DOI:
10.1145/2505515.2505635
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Michael R. Genesereth
中科院分区:
文献类型:
--
作者:
Sudhir Agarwal;Michael R. Genesereth
For increasingly sophisticated use cases end users often need to extract, combine, and aggregate information from various (often dynamically generated) web pages from multiple websites. Current search engines do not focus on combining information from various web pages in order to answer the overall information need of the user. Semantic Web and Linked Data usually take a static view on the data and rely on providers' cooperation. In this paper, we present a novel approach that enables end users to easily extract data from web pages while they browse, store it locally in their browser as well as structure, integrate and search such data. We propose Datalog rules for integrating and searching the extracted data. We show how cleaning steps and integration rules can be reused to accelerate the cleaning and integration of extracted data. The proposed approach is implemented as a browser plugin. We present its implementation details and report on our evaluation of the plugin concerning user experience and browsing time saving.