Cross-Language Record Linkage based on Semantic Matching of Metadata
Cross-Language Record Linkage based on Semantic Matching of Metadata
复制标题
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Akira Maeda
中科院分区:
文献类型:
--
作者:
Akira Maeda
Record linkage is finding record pairs that refer to the same entities or objects across multiple data sources. This task is crucial in various research fields, such as federated search and data integration. This article focuses on the new challenge of cross-language record linkage, where records are from data sources in different languages. To compare the records in different languages, the records’ metadata needs to be translated from a source language to a target language. This causes mismatches between the translated metadata and the metadata in the target language, since similar meanings are sometimes expressed by different words during translation. Thus, conventional string-based similarity metrics are insufficient for measuring the similarities between the translated metadata and the metadata in the target language. Therefore, we propose a method of dealing with the mismatching problem in cross-language record linkage. For each translated metadata of the source language, first, we use a string-based similarity metric to identify the potential matching metadata in the target language as candidates. Then, we employ word embeddings to perform the semantic matching between the translated metadata and its candidate matching metadata in the target language. Our method is evaluated on a real-world dataset in Japanese and English. Our experiments proved that our proposed method outperforms baseline methods that only rely on string similarities or the semantic matching method.