Revisiting Conditional Functional Dependency Discovery: Splitting the "C" from the "FD"

Revisiting Conditional Functional Dependency Discovery: Splitting the "C" from the "FD"
复制标题

DOI:
10.1007/978-3-030-10928-8_33
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Joeri Rammelaere;Floris Geerts
Joeri Rammelaere;Floris Geerts
中科院分区:
其他
文献类型:
--
作者:
Joeri Rammelaere;Floris Geerts

文献摘要

被引文献

相似文献

清理脏数据的许多技术都是基于实施某些完整性约束集的。条件函数依赖(CFD)是传统函数依赖(FDs)和关联规则的结合,被广泛用作数据清洗的约束形式。然而,这类差价合约的发现受到的关注有限。本文将CFD看作是关联规则的扩展,提出了三种通用的(近似)CFD发现方法,每种方法都使用了一种不同的方式来结合模式挖掘来发现条件(CFD中的C)和FD发现。我们讨论了现有算法如何适应这三种方法,并引入了新的技术来改进发现过程。结果表明,与传统的CFD发现方法CTane相比,选择正确的方法可以提高性能。与本文相关的代码请访问:https://github.com/j-r77/cfddiscovery,https://codeocean.com/2018/06/20/discovering-conditional-functional-dependencies/code。
Many techniques for cleaning dirty data are based on enforcing some set of integrity constraints. Conditional functional dependencies (CFDs) are a combination of traditional Functional dependencies (FDs) and association rules, and are widely used as a constraint formalism for data cleaning. However, the discovery of such CFDs has received limited attention. In this paper, we regard CFDs as an extension of association rules, and present three general methodologies for (approximate) CFD discovery, each using a different way of combining pattern mining for discovering the conditions (the “C” in CFD) with FD discovery. We discuss how existing algorithms fit into these three methodologies, and introduce new techniques to improve the discovery process. We show that the right choice of methodology improves performance over the traditional CFD discovery method CTane. Code related to this paper is available at: https://github.com/j-r77/cfddiscovery , https://codeocean.com/2018/06/20/discovering-conditional-functional-dependencies/code .