Inferring Variable Labels Considering Co-occurrence of Variable Labels in Data Jackets
Inferring Variable Labels Considering Co-occurrence of Variable Labels in Data Jackets
复制标题
DOI:
10.1109/icdmw.2016.0115
复制
发表时间:
2016-12
期刊:
影响因子:
--
通讯作者:
Teruaki Hayashi
中科院分区:
文献类型:
--
作者:
Teruaki Hayashi
Data Jacket (DJ) is a technique for sharing information about data and for considering the potential value of datasets, allowing data itself hidden, by describing the summary of data in natural language. In DJs, variables in datasets are described as variable labels (VLs), which is the name/meaning of variables. In the context of data utilization and exchange, the utility of data can be discussed upon the VLs to consider the combination of data stored in different domains. However, due to the lack of VLs in some DJs, DJs essentially related to each other cannot be formed to have linkage through string matching of VLs, which makes it difficult to think of feasible plans of data analyses and combinations. In this paper, we propose a method for inferring VLs in DJs whose VLs are missing or unknown, using the content of the outlines of DJs written in free texts. Specifically, we focus on the co-occurrence of VLs in DJs. The co-occurrence of VLs is a feature that there may be a highly frequent pair of VLs appearing at the same time, e.g., "year" and "day", or "name" and "gender". By focusing on the co-occurrence of VLs and the similarity of the outlines of data in training data of DJs, we demonstrate that our proposed method works significantly better than the method introducing only the similarity of outlines of datasets.