Inferring variable labels using outlines of data in Data Jackets by considering similarity and co-occurrence
Inferring variable labels using outlines of data in Data Jackets by considering similarity and co-occurrence
复制标题
DOI:
10.1007/s41060-018-0152-8
复制
发表时间:
2018-09
影响因子:
2.4
通讯作者:
Teruaki Hayashi;Y. Ohsawa
中科院分区:
文献类型:
--
作者:
Teruaki Hayashi;Y. Ohsawa
The Data Jacket (DJ) is a technique for sharing information related to data, where the data are hidden, by summarizing them in natural language. In DJs, variables are described by variable labels (VLs), which are the names/meanings of variables, and the utility of data is estimated through combinations of VLs. However, DJs do not always contain VLs because the rules describing DJs cannot compel data owners to enter all relevant information. Owing to a lack of VLs in some DJs, even if the DJs can be combined, their combinations cannot be implemented through the string matching of the VLs. In this paper, we propose a method for inferring VLs in DJs using the text in their outlines. We focus on similarity among the outlines of DJs and create two models for inferring VLs, i.e., based on the similarity of the outlines and the co-occurrence of the VLs. We implemented our models on a similarity and a co-occurrence matrix and applied the proposed method to two types of test data: the DJs of public data and business data. The results of experiments show that our method is significantly superior to the technique that uses only the string matching of the VLs.