Inferring variable labels using outlines of data in Data Jackets by considering similarity and co-occurrence

Inferring variable labels using outlines of data in Data Jackets by considering similarity and co-occurrence
复制标题

DOI:
10.1007/s41060-018-0152-8
复制
发表时间:
2018-09
影响因子:
2.4
通讯作者:
Teruaki Hayashi;Y. Ohsawa
Teruaki Hayashi;Y. Ohsawa
中科院分区:
--
文献类型:
--
作者:
Teruaki Hayashi;Y. Ohsawa

文献摘要

相似文献

数据夹(DJ)是一种用于共享与数据相关的信息的技术,其中数据是隐藏的,通过用自然语言总结它们。在DJ中,变量由变量标签(VL)描述,变量标签是变量的名称/含义,并且通过VL的组合来估计数据的效用。然而,DJ并不总是包含VL,因为描述DJ的规则不能强制数据所有者输入所有相关信息。由于某些DJ中缺少VL,即使DJ可以组合,也无法通过VL的字符串匹配来实现它们的组合。在本文中,我们提出了一种方法来推断在DJ的VL使用的文本在他们的大纲。我们专注于DJ的轮廓之间的相似性,并创建两个模型来推断VL,即,基于轮廓的相似性和VL的共同出现。我们实现了我们的模型上的相似性和共生矩阵,并将所提出的方法应用到两种类型的测试数据:公共数据和商业数据的DJ。实验结果表明,该方法明显优于仅使用字符串匹配的方法,具有上级的优越性。
The Data Jacket (DJ) is a technique for sharing information related to data, where the data are hidden, by summarizing them in natural language. In DJs, variables are described by variable labels (VLs), which are the names/meanings of variables, and the utility of data is estimated through combinations of VLs. However, DJs do not always contain VLs because the rules describing DJs cannot compel data owners to enter all relevant information. Owing to a lack of VLs in some DJs, even if the DJs can be combined, their combinations cannot be implemented through the string matching of the VLs. In this paper, we propose a method for inferring VLs in DJs using the text in their outlines. We focus on similarity among the outlines of DJs and create two models for inferring VLs, i.e., based on the similarity of the outlines and the co-occurrence of the VLs. We implemented our models on a similarity and a co-occurrence matrix and applied the proposed method to two types of test data: the DJs of public data and business data. The results of experiments show that our method is significantly superior to the technique that uses only the string matching of the VLs.