An Interactive Data Quality Test Approach for Constraint Discovery and Fault Detection
An Interactive Data Quality Test Approach for Constraint Discovery and Fault Detection
复制标题
DOI:
10.1109/bigdata47090.2019.9006446
复制
发表时间:
2019-12
期刊:
影响因子:
--
通讯作者:
Hajar Homayouni;Sudipto Ghosh;I. Ray;M. Kahn
中科院分区:
文献类型:
--
作者:
Hajar Homayouni;Sudipto Ghosh;I. Ray;M. Kahn
Data quality tests validate heterogeneous data to detect violations of syntactic and semantic constraints. The specification of these constraints can be incomplete because domain experts typically specify them in an ad hoc manner. Existing automated test approaches can generate false alarms and do not explain the constraint violations while reporting faulty data records. In previous work, we proposed ADQuaTe, which is an automated data quality test approach that uses an unsupervised deep learning techni que (1) to discover constraints from big datasets that may have been missed by experts, and (2) to label as suspicious those records that violate the constraints. These records are grouped and explanations for constraint violations are presented to domain experts who determine whether or not the groups are actually faulty. This paper presents ADQuaTe2, which extends ADQuaTe to use an interactive learning technique that incorporates expert feedback to retrain the learning model and improve the accuracy of constraint discovery and fault detection. We evaluate the effectiveness of the approach on real-world datasets from a health data warehouse and a plant diagnosis database. We also use datasets with known faults from the UCI repository to evaluate the improvement in the accuracy of the approach after incorporating ground truth knowledge.