An Interactive Data Quality Test Approach for Constraint Discovery and Fault Detection

An Interactive Data Quality Test Approach for Constraint Discovery and Fault Detection
复制标题

DOI:
10.1109/bigdata47090.2019.9006446
复制
发表时间:
2019-12
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Hajar Homayouni;Sudipto Ghosh;I. Ray;M. Kahn
Hajar Homayouni;Sudipto Ghosh;I. Ray;M. Kahn
中科院分区:
其他
文献类型:
--
作者:
Hajar Homayouni;Sudipto Ghosh;I. Ray;M. Kahn

文献摘要

相似文献

数据质量测试验证异类数据,以检测违反语法和语义约束的情况。这些约束的规范可能是不完整的,因为领域专家通常以特别的方式指定它们。现有的自动化测试方法会产生错误警报,并且在报告错误数据记录时不能解释违反约束的情况。在以前的工作中,我们提出了ADQUATE,这是一种自动化的数据质量测试方法,它使用无监督的深度学习技术(1)从大数据集中发现可能被专家遗漏的约束,以及(2)将违反约束的记录标记为可疑。对这些记录进行分组,并将违反约束的解释提供给领域专家,他们确定这些组是否确实有问题。本文提出了ADQuaTe2,它对ADQUATE进行了扩展,使用了结合专家反馈的交互式学习技术来重新训练学习模型,提高了约束发现和故障检测的准确性。我们在一个健康数据仓库和一个植物诊断数据库的真实数据集上评估了该方法的有效性。我们还使用来自UCI储存库的具有已知故障的数据集来评估在纳入基本事实知识后该方法在准确性方面的改进。
Data quality tests validate heterogeneous data to detect violations of syntactic and semantic constraints. The specification of these constraints can be incomplete because domain experts typically specify them in an ad hoc manner. Existing automated test approaches can generate false alarms and do not explain the constraint violations while reporting faulty data records. In previous work, we proposed ADQuaTe, which is an automated data quality test approach that uses an unsupervised deep learning techni que (1) to discover constraints from big datasets that may have been missed by experts, and (2) to label as suspicious those records that violate the constraints. These records are grouped and explanations for constraint violations are presented to domain experts who determine whether or not the groups are actually faulty. This paper presents ADQuaTe2, which extends ADQuaTe to use an interactive learning technique that incorporates expert feedback to retrain the learning model and improve the accuracy of constraint discovery and fault detection. We evaluate the effectiveness of the approach on real-world datasets from a health data warehouse and a plant diagnosis database. We also use datasets with known faults from the UCI repository to evaluate the improvement in the accuracy of the approach after incorporating ground truth knowledge.