Statistical Distortion: Consequences of Data Cleaning

Statistical Distortion: Consequences of Data Cleaning
复制标题

DOI:
10.14778/2350229.2350279
复制
发表时间:
2012-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
T. Dasu;J. Loh
T. Dasu;J. Loh
中科院分区:
其他
文献类型:
--
作者:
T. Dasu;J. Loh

文献摘要

被引文献

相似文献

我们引入统计失真的概念,作为衡量数据清洗策略有效性的重要指标。我们使用这个指标,提出了一个广泛适用的,但可扩展的实验框架,沿沿着三个维度:故障改善,统计失真和成本相关的标准评估数据清洗策略。现有的指标关注故障改善和成本,但不关注数据清理策略的统计影响。我们用一套全面的实验和分析来说明我们在真实的世界数据上的框架。
We introduce the notion of statistical distortion as an essential metric for measuring the effectiveness of data cleaning strategies. We use this metric to propose a widely applicable yet scalable experimental framework for evaluating data cleaning strategies along three dimensions: glitch improvement, statistical distortion and cost-related criteria. Existing metrics focus on glitch improvement and cost, but not on the statistical impact of data cleaning strategies. We illustrate our framework on real world data, with a comprehensive suite of experiments and analyses.