Extraction and integration of web data by end-users

Extraction and integration of web data by end-users
复制标题

最终用户提取和整合网络数据

DOI:
10.1145/2505515.2505635
复制
发表时间:
2013
期刊:
Proceedings of the 22nd ACM international conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Michael R. Genesereth
Michael R. Genesereth
中科院分区:
--
文献类型:
--
作者:
Sudhir Agarwal;Michael R. Genesereth

文献摘要

被引文献

相似文献

对于日益复杂的用例,最终用户通常需要从多个网站的各种(通常是动态生成的)网页中提取、组合和聚合信息。当前的搜索引擎并不专注于组合来自各个网页的信息来满足用户的总体信息需求。语义网和链接数据通常采用静态数据视图并依赖于提供商的合作。在本文中,我们提出了一种新颖的方法,使最终用户能够在浏览时轻松地从网页中提取数据,将其本地存储在浏览器中,以及构建、集成和搜索此类数据。我们提出了用于集成和搜索提取的数据的数据记录规则。我们展示了如何重用清理步骤和集成规则来加速提取数据的清理和集成。所提出的方法是作为浏览器插件实现的。我们介绍其实施细节,并报告我们对该插件在用户体验和浏览时间节省方面的评估。
For increasingly sophisticated use cases end users often need to extract, combine, and aggregate information from various (often dynamically generated) web pages from multiple websites. Current search engines do not focus on combining information from various web pages in order to answer the overall information need of the user. Semantic Web and Linked Data usually take a static view on the data and rely on providers' cooperation. In this paper, we present a novel approach that enables end users to easily extract data from web pages while they browse, store it locally in their browser as well as structure, integrate and search such data. We propose Datalog rules for integrating and searching the extracted data. We show how cleaning steps and integration rules can be reused to accelerate the cleaning and integration of extracted data. The proposed approach is implemented as a browser plugin. We present its implementation details and report on our evaluation of the plugin concerning user experience and browsing time saving.