DeepClean: Data Cleaning via Question Asking

DeepClean: Data Cleaning via Question Asking
复制标题

DOI:
10.1109/dsaa.2018.00039
复制
发表时间:
2018-10
期刊:
2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA)
影响因子:
--
通讯作者:
Xinyang Zhang-;Yujie Ji;Chanh Nguyen;Ting Wang
Xinyang Zhang-;Yujie Ji;Chanh Nguyen;Ting Wang
中科院分区:
其他
文献类型:
--
作者:
Xinyang Zhang-;Yujie Ji;Chanh Nguyen;Ting Wang

文献摘要

相似文献

作为数据分析流程中的一项关键任务,数据清理是众所周知的人力密集型任务并且容易出错。知识库辅助数据清理已被证明是发现和修复数据缺陷的强大工具;然而,它的适用性不可避免地受到知识库的自然局限性的限制。与此同时,尽管大量知识源以自由文本语料库(例如维基百科)的形式存在,但将它们转换为现有数据清理工具可用的格式可能成本高昂且容易出错,甚至根本不可能。在这里,我们介绍 DeepClean,第一个由自由文本知识源支持的端到端数据清理框架。在较高层面上,DeepClean 通过其问答 (QA) 界面利用知识源,并通过迭代提问实现高质量的清洁。具体来说,DeepClean分三个阶段检测和修复数据缺陷:(i)模式提取——自动发现数据属性的语义类型及其相关性; (ii) 问题生成——它将每个数据元组转换为最小的验证问题集; (iii) 完成和修复 - 通过对照数据值检查知识源返回的答案,识别错误情况并提出可能的修复建议。通过广泛的实证研究,我们证明 DeepClean 适用于一系列领域,并且可以有效修复各种数据缺陷,强调自由文本知识源驱动的数据清理是未来研究的一个有前途的方向。
As one critical task in the data analysis pipeline, data cleaning is notoriously human labor-intensive and error-prone. Knowledge base-assisted data cleaning has proved a powerful tool for finding and fixing data defects; however, its applicability is inevitably bounded by the natural limitations of knowledge bases. Meanwhile, although a vast number of knowledge sources exist in the form of free-text corpora (e.g., Wikipedia), transforming them into formats usable by existing data cleaning tools can be prohibitively costly and error-prone, if not at all impossible. Here, we present DeepClean, the first end-to-end data cleaning framework powered by free-text knowledge sources. At a high level, DeepClean leverages a knowledge source through its question-answering (QA) interface and achieves high-quality cleaning via iterative question asking. Specifically, DeepClean detects and repairs data defects in three stages: (i) Pattern extraction - it automatically discovers the semantic types of the data attributes as well as their correlations; (ii) Question generation - it translates each data tuple into a minimal set of validation questions; (iii) Completion and repair - by checking the answers returned by the knowledge source against the data values, it identifies erroneous cases and suggests possible fixes. Through extensive empirical studies, we demonstrate that DeepClean is applicable to a range of domains, and can effectively repair a variety of data defects, highlighting data cleaning powered by free-text knowledge sources as a promising direction for future research.