Top-k generation of integrated schemas based on directed and weighted correspondences

Top-k generation of integrated schemas based on directed and weighted correspondences
复制标题

DOI:
10.1145/1559845.1559913
复制
发表时间:
2009-06
期刊:
Proceedings of the 2009 ACM SIGMOD International Conference on Management of data
影响因子:
--
通讯作者:
A. Radwan;Lucian Popa;I. Stanoi;A. Younis
A. Radwan;Lucian Popa;I. Stanoi;A. Younis
中科院分区:
其他
文献类型:
--
作者:
A. Radwan;Lucian Popa;I. Stanoi;A. Younis

文献摘要

被引文献

相似文献

模式集成是基于一组现有的源模式和一组匹配源模式的对应关系创建统一目标模式的问题。以前的模式集成方法依赖于对集成模式可能的多种设计选择的隐式或显式探索。这种探索严重依赖于用户交互;因此,它是耗时且劳动密集型的。此外,先前的方法忽略了通常由模式匹配过程产生的附加信息,即,权重,并且在某些情况下,忽略了与对应性相关联的方向。在本文中,我们提出了一个更自动化的方法,模式集成的基础上使用的定向和加权之间的对应关系,出现在源模式的概念。我们的方法的一个关键组成部分是一个新的top-k排名算法自动生成的最佳候选模式。该算法赋予模式更多的权重,联合收割机组合的概念具有较高的相似性或覆盖率。因此,该算法做出了某些决策,否则这些决策可能会由人类专家做出。我们表明,该算法在多项式时间内运行,而且在实践中具有良好的性能。
Schema integration is the problem of creating a unified target schema based on a set of existing source schemas and based on a set of correspondences that are the result of matching the source schemas. Previous methods for schema integration rely on the exploration, implicit or explicit, of the multiple design choices that are possible for the integrated schema. Such exploration relies heavily on user interaction; thus, it is time consuming and labor intensive. Furthermore, previous methods have ignored the additional information that typically results from the schema matching process, that is, the weights and in some cases the directions that are associated with the correspondences. In this paper, we propose a more automatic approach to schema integration that is based on the use of directed and weighted correspondences between the concepts that appear in the source schemas. A key component of our approach is a novel top-k ranking algorithm for the automatic generation of the best candidate schemas. The algorithm gives more weight to schemas that combine the concepts with higher similarity or coverage. Thus, the algorithm makes certain decisions that otherwise would likely be taken by a human expert. We show that the algorithm runs in polynomial time and moreover has good performance in practice.