Cross-Language Record Linkage based on Semantic Matching of Metadata

Cross-Language Record Linkage based on Semantic Matching of Metadata
复制标题

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Akira Maeda
Akira Maeda
中科院分区:
其他
文献类型:
--
作者:
Akira Maeda

文献摘要

相似文献

记录链接是在多个数据源中查找引用相同实体或对象的记录对。这项任务在许多研究领域都是至关重要的,例如联邦搜索和数据集成。本文主要关注跨语言记录链接的新挑战,其中记录来自不同语言的数据源。为了比较不同语言的记录,需要将记录的元数据从源语言翻译成目标语言。这将导致翻译后的元数据与目标语言中的元数据之间的不匹配,因为在翻译过程中,相似的含义有时由不同的单词表达。因此,传统的基于字符串的相似性度量不足以度量翻译后的元数据与目标语言的元数据之间的相似性。因此,我们提出了一种处理跨语言记录链接中的不匹配问题的方法。对于源语言的每个翻译元数据,首先,我们使用基于字符串的相似性度量来识别目标语言中潜在的匹配元数据作为候选。然后,我们使用词嵌入在翻译后的元数据和目标语言的候选匹配元数据之间进行语义匹配。我们的方法在日语和英语的真实数据集上进行了评估。实验证明,本文提出的方法优于仅依赖字符串相似度或语义匹配方法的基线方法。
Record linkage is finding record pairs that refer to the same entities or objects across multiple data sources. This task is crucial in various research fields, such as federated search and data integration. This article focuses on the new challenge of cross-language record linkage, where records are from data sources in different languages. To compare the records in different languages, the records’ metadata needs to be translated from a source language to a target language. This causes mismatches between the translated metadata and the metadata in the target language, since similar meanings are sometimes expressed by different words during translation. Thus, conventional string-based similarity metrics are insufficient for measuring the similarities between the translated metadata and the metadata in the target language. Therefore, we propose a method of dealing with the mismatching problem in cross-language record linkage. For each translated metadata of the source language, first, we use a string-based similarity metric to identify the potential matching metadata in the target language as candidates. Then, we employ word embeddings to perform the semantic matching between the translated metadata and its candidate matching metadata in the target language. Our method is evaluated on a real-world dataset in Japanese and English. Our experiments proved that our proposed method outperforms baseline methods that only rely on string similarities or the semantic matching method.