Incrementally improving dataspaces based on user feedback

Incrementally improving dataspaces based on user feedback
复制标题

DOI:
10.1016/j.is.2013.01.006
复制
发表时间:
2013-07-01
影响因子:
3.7
通讯作者:
Hedeler, Cornelia
Hedeler, Cornelia
中科院分区:
计算机科学2区
文献类型:
--
作者:
Belhajjame, Khalid;Paton, Norman W.;Hedeler, Cornelia

文献摘要

被引文献

相似文献

数据空间愿景的一个方面被阐明为提供传统数据集成的各种好处,并降低前期成本。在本文中,我们介绍了旨在通过以即付即用的方式与最终用户交互来支持模式映射规范的技术。特别是,我们展示了如何使用现有的匹配和映射生成技术自动获得的模式映射,可以使用从最终用户获得的查询结果的反馈来评估其对用户需求的适合度的度量来进行标注,使用基于用户反馈计算的标注,并且在给定用户在查准率和召回率方面的要求的情况下,我们提出了一种选择产生满足所述需求的结果的映射集的方法。在这样做的过程中,我们将映射选择视为一个优化问题。反馈可能会显示模式映射的质量很差。我们展示了如何使用映射注释来支持通过精化从现有映射派生出质量更好的映射。进化算法被用来高效地探索可以通过精化获得的大量映射空间,用户反馈也可以用来标注用户针对集成模式提出的查询的结果。我们展示了如何为这样的查询计算精度和召回率的估计。我们还研究了将关于(集成)查询结果的反馈向下传播到用于填充集成模式中的基本关系的映射的问题。(C)2013爱思唯尔有限公司。保留所有权利。
One aspect of the vision of dataspaces has been articulated as providing various benefits of classical data integration with reduced up-front costs. In this paper, we present techniques that aim to support schema mapping specification through interaction with end users in a pay-as-you-go fashion. In particular, we show how schema mappings, that are obtained automatically using existing matching and mapping generation techniques, can be annotated with metrics estimating their fitness to user requirements using feedback on query results obtained from end users.Using the annotations computed on the basis of user feedback, and given user requirements in terms of precision and recall, we present a method for selecting the set of mappings that produce results meeting the stated requirements. In doing so, we cast mapping selection as an optimization problem. Feedback may reveal that the quality of schema mappings is poor. We show how mapping annotations can be used to support the derivation of better quality mappings from existing mappings through refinement. An evolutionary algorithm is used to efficiently and effectively explore the large space of mappings that can be obtained through refinement.User feedback can also be used to annotate the results of the queries that the user poses against an integration schema. We show how estimates for precision and recall can be computed for such queries. We also investigate the problem of propagating feedback about the results of (integration) queries down to the mappings used to populate the base relations in the integration schema. (C) 2013 Elsevier Ltd. All rights reserved.