Efficient Web-Based Data Imputation with Graph Model

Efficient Web-Based Data Imputation with Graph Model
复制标题

DOI:
10.1007/978-3-319-55705-2_17
复制
发表时间:
2016-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Yiwen Tang;Hongzhi Wang;Shiwei Zhang;Huijun Zhang;Ruoxi Shi
Yiwen Tang;Hongzhi Wang;Shiwei Zhang;Huijun Zhang;Ruoxi Shi
中科院分区:
其他
文献类型:
--
作者:
Yiwen Tang;Hongzhi Wang;Shiwei Zhang;Huijun Zhang;Ruoxi Shi

文献摘要

被引文献

相似文献

数据估算的一个挑战是缺乏知识。在本文中,我们试图解决这个挑战,涉及额外的知识,从网络。为了实现高性能的基于Web的插补,我们使用依赖性,即FD和CFD,自动插补尽可能多的值,并以最少的Web访问来填充其他缺失值,其成本相对较大。为了充分利用依赖关系,我们将数据上的依赖关系集建模为图,并基于这种图模型执行基于Web的自动填充和关键字生成。利用生成的关键字,我们设计了两个算法来从搜索结果中提取用于填补的值。基于真实数据集的大量实验结果表明,与现有方法相比,该方法可以有效地填补缺失值。
A challenge for data imputation is the lack of knowledge. In this paper, we attempt to address this challenge by involving extra knowledge from web. To achieve high-performance web-based imputation, we use the dependency, i.e. FDs and CFDs, to impute as many as possible values automatically and fill in the other missing values with the minimal access of web, whose cost is relatively large. To make sufficient use of dependencies, we model the dependency set on the data as a graph and perform automatical imputation and keywords generation for web-based imputation based on such graph model. With the generated keywords, we design two algorithms to extract values for imputation from the search results. Extensive experimental results based on real-world data collections show that the proposed approach could impute missing values efficiently and effectively compared to existing approach.