Cross-Language Record Linkage by Exploiting Semantic Matching of Textual Metadata

Cross-Language Record Linkage by Exploiting Semantic Matching of Textual Metadata
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Yuting Song;Taisuke Kimura;B. Batjargal;Akira Maeda
Yuting Song;Taisuke Kimura;B. Batjargal;Akira Maeda
中科院分区:
其他
文献类型:
--
作者:
Yuting Song;Taisuke Kimura;B. Batjargal;Akira Maeda

文献摘要

相似文献

Cross-language record linkage is a task of finding pairs of records that refer to the same entity across multiple databases in different languages. It is crucial to various research fields, such as federated search and data integration. The matching of textual values of metadata fields plays an important part in comparing record pairs. When matching textual values across languages, one problem is that the mismatches between semantically related translations of metadata values in source language and metadata values in target language, which refer to the same entity. For example, when comparing the records in Japanese (source language) and English (target language), the Japanese word “白雨” in metadata is translated into “rainfall”. However, the corresponding word in English metadata is “storm”, which is semantically related to “rainfall”. As a consequence, the commonly used string-based matching cannot measure the relevance of semantically related words. In this paper, we propose a method for semantic matching of textual metadata, which is based on word embedding that can capture the semantic similarity relationships among words. The effectiveness of this method is evaluated on film related textual metadata in Japanese and English. Then, we use our method to link the identical Ukiyo-e prints between the databases in Japanese and English.