Multi-source relational data fusion

Multi-source relational data fusion
复制标题

DOI:
10.1360/ssi-2019-0172
复制
发表时间:
2020-04
影响因子:
--
通讯作者:
Yue Ding;Juan Wang;Wei Lu;Chuitian Rong;Xiaoyong Du
Yue Ding;Juan Wang;Wei Lu;Chuitian Rong;Xiaoyong Du
中科院分区:
--
文献类型:
--
作者:
Yue Ding;Juan Wang;Wei Lu;Chuitian Rong;Xiaoyong Du

文献摘要

相似文献

针对“信息孤岛”环境下的关系数据融合问题,提出了一种多源关系数据融合框架。该框架由三个部分组成:模式匹配,实体对齐和实体融合。基于匈牙利算法,提出了一种多源关系数据属性对齐发现机制。通过提取属性值的多维特征,有效地实现了多源关系数据的模式匹配。为了连接多源数据的元组对,我们引入了多样性抽样策略和实体特征提取方法。这些可以有效地提高实体对齐的性能。最后,将链接的实体进行融合,以提供数据分析的统一视图。为了验证所提出的方法的有效性和效率,我们实现了一个融合系统称为数据拼图,这是验证与真实的公共多领域的数据。实验结果表明,该方法能有效地融合多源关系数据,具有较高的查全率和查准率。
Focusing on the problem of relational data fusion in the environment with “information isolated island”, this paper presents a multi-sources relational data fusion (MSF) framework. The framework consists of three components: schema matching, entity alignment, and entity fusion. Based on the Hungarian algorithm, we propose an alignment discovery mechanism for the attributes alignment among multi-sources relational data. By extracting the multi-dimensional features of attribute values, we efficiently realized schema matching of multi-sources relational data. To link the tuple pairs from multi-source data, we introduced the diversity sampling strategy and the entity feature extraction approach. These can effectively improve the performance of entity alignment. Finally, linked entities are fused to provide a unified view of data analysis. To verify the usefulness and efficiency of the proposed methods, we implemented a fusion system called Data Puzzle, which is verified with the real public multi-field data. Experimental results demonstrate that the proposed methods can fuse multi-source relational data efficiently with high recall and precision.