Integrating web query results: holistic schema matching

Integrating web query results: holistic schema matching
复制标题

DOI:
10.1145/1458082.1458090
复制
发表时间:
2008-10
期刊:
--
影响因子:
--
通讯作者:
Shui-Lung Chuang;K. Chang
Shui-Lung Chuang;K. Chang
中科院分区:
其他
文献类型:
--
作者:
Shui-Lung Chuang;K. Chang

文献摘要

被引文献

相似文献

大量在线数据源的出现迫切需要更自动但更准确的数据集成技术。对于查询这些数据源返回的数据,大多数工作都集中在如何更准确地提取嵌入的结构化数据。然而,为了最终提供对这些查询结果的集成访问,最后但并非最不重要的步骤是联合收割机组合来自不同源的提取数据。一个关键的任务是找到源之间数据字段的对应关系-一个众所周知的模式匹配问题。查询结果是一个小的和有偏见的样本集的实例从源获得的模式信息,因此是非常隐式和不完整的,这往往会妨碍现有的模式匹配方法有效地执行。在本文中,我们开发了一种新的框架,用于理解和有效地支持模式匹配等基于实例的数据,特别是集成多个源。我们将发现匹配视为构建最能描述输入数据的更完整的域模式。有了这个概念视图,我们可以利用各种数据实例,并通过整体的、多源的模式匹配无缝地观察到数据,以获得更准确的匹配结果。我们的实验表明,我们的框架始终优于基线成对和基于聚类的方法(提高F-测量从50-89%到89-94%),并一致适用于调查领域。
The emergence of numerous data sources online has presented a pressing need for more automatic yet accurate data integration techniques. For the data returned from querying such sources, most works focus on how to extract the embedded structured data more accurately. However, to eventually provide an integrated access to these query results, a last but not least step is to combine the extracted data coming from different sources. A critical task is finding the correspondence of the data fields between the sources - a problem well known as schema matching. Query results are a small and biased sample set of instances obtained from sources; the obtained schema information is thus very implicit and incomplete, which often prevents existing schema matching approaches from performing effectively. In this paper, we develop a novel framework for understanding and effectively supporting schema matching on such instance-based data, especially for integrating multiple sources. We view discovering matching as constructing a more complete domain schema that best describes the input data. With this conceptual view, we can leverage various data instances and observed regularities seamlessly with holistic, multiple-source schema matching to achieve more accurate matching results. Our experiments show that our framework consistently outperforms baseline pairwise and clustering-based approaches (raising F-measure from 50-89% to 89-94%) and works uniformly well for the surveyed domains.