Inferring Variable Labels Considering Co-occurrence of Variable Labels in Data Jackets

Inferring Variable Labels Considering Co-occurrence of Variable Labels in Data Jackets
复制标题

DOI:
10.1109/icdmw.2016.0115
复制
发表时间:
2016-12
期刊:
2016 IEEE 16th International Conference on Data Mining Workshops (ICDMW)
影响因子:
--
通讯作者:
Teruaki Hayashi
Teruaki Hayashi
中科院分区:
其他
文献类型:
--
作者:
Teruaki Hayashi

文献摘要

相似文献

数据夹(DJ)是一种技术,用于共享有关数据的信息,并考虑数据集的潜在价值,允许数据本身隐藏,通过用自然语言描述数据的摘要。在DJ中,数据集中的变量被描述为变量标签(VL),这是变量的名称/含义。在数据利用和交换的上下文中,可以在虚拟L上讨论数据的效用,以考虑存储在不同域中的数据的组合。然而,由于某些DJ中缺少VL,因此无法通过VL的字符串匹配来形成本质上相互关联的DJ具有链接,这使得难以想到可行的数据分析和组合方案。在本文中,我们提出了一种方法来推断的虚拟现实中的DJ的虚拟现实缺失或未知,使用的内容的自由文本中写的DJ的大纲。具体来说,我们专注于在DJ的VL的共同出现。VL的共现是可能存在同时出现的高频率的VL对的特征,例如,“年”和“日”,或“姓名”和“性别”。通过重点放在VLs的同现和DJ的训练数据中的数据的轮廓的相似性,我们证明了我们提出的方法的效果显着优于只引入数据集的轮廓的相似性的方法。
Data Jacket (DJ) is a technique for sharing information about data and for considering the potential value of datasets, allowing data itself hidden, by describing the summary of data in natural language. In DJs, variables in datasets are described as variable labels (VLs), which is the name/meaning of variables. In the context of data utilization and exchange, the utility of data can be discussed upon the VLs to consider the combination of data stored in different domains. However, due to the lack of VLs in some DJs, DJs essentially related to each other cannot be formed to have linkage through string matching of VLs, which makes it difficult to think of feasible plans of data analyses and combinations. In this paper, we propose a method for inferring VLs in DJs whose VLs are missing or unknown, using the content of the outlines of DJs written in free texts. Specifically, we focus on the co-occurrence of VLs in DJs. The co-occurrence of VLs is a feature that there may be a highly frequent pair of VLs appearing at the same time, e.g., "year" and "day", or "name" and "gender". By focusing on the co-occurrence of VLs and the similarity of the outlines of data in training data of DJs, we demonstrate that our proposed method works significantly better than the method introducing only the similarity of outlines of datasets.