A Fast Linkage Detection Scheme for Multi-Source Information Integration

A Fast Linkage Detection Scheme for Multi-Source Information Integration
复制标题

DOI:
10.1109/wiri.2005.2
复制
发表时间:
2005-04
期刊:
International Workshop on Challenges in Web Information Retrieval and Integration
影响因子:
--
通讯作者:
Akiko Aizawa;K. Oyama
Akiko Aizawa;K. Oyama
中科院分区:
其他
文献类型:
--
作者:
Akiko Aizawa;K. Oyama

文献摘要

相似文献

记录链接是指识别与相同现实世界实体关联的记录的技术。记录链接不仅对于集成独立生成的多源数据库至关重要,而且也被认为是集成异构Web资源的关键问题之一。然而,当针对大规模数据时,枚举所有可能的联系的成本通常变得高得不切实际。基于这样的背景,本文提出了一种快速高效的连锁检测方法。该方法的特点是:首先,它利用了后缀数组结构,可以使用可变长度的 n 元语法进行链接检测。其次,它使用从已知的可靠链接中提取的“阻塞键”动态生成可能关联的记录块。本文还报告了我们初步实验的结果,其中所提出的方法应用于四个书目数据库的集成,这些数据库的规模扩大到超过 1000 万条记录。
Record linkage refers to techniques for identifying records associated with the same real-world entities. Record linkage is not only crucial in integrating multi-source databases that have been generated independently, but is also considered to be one of the key issues in integrating heterogeneous Web resources. However, when targeting large-scale data, the cost of enumerating all the possible linkages often becomes impracticably high. Based on this background, this paper proposes a fast and efficient method for linkage detection. The features of the proposed approach are: first, it exploits a suffix array structure that enables linkage detection using variable length n-grams. Second, it dynamically generates blocks of possibly associated records using ‘blocking keys’ extracted from already known reliable linkages. The results from our preliminary experiments where the proposed method was applied to the integration of four bibliographic databases, which scale up to more than 10 million records, are also reported in the paper.