Conformance Constraint Discovery: Measuring Trust in Data-Driven Systems

Conformance Constraint Discovery: Measuring Trust in Data-Driven Systems
复制标题

DOI:
10.1145/3448016.3452795
复制
发表时间:
2021-06
期刊:
Proceedings of the 2021 International Conference on Management of Data
影响因子:
--
通讯作者:
Anna Fariha;A. Tiwari;Arjun Radhakrishna;Sumit Gulwani;A. Meliou
Anna Fariha;A. Tiwari;Arjun Radhakrishna;Sumit Gulwani;A. Meliou
中科院分区:
其他
文献类型:
--
作者:
Anna Fariha;A. Tiwari;Arjun Radhakrishna;Sumit Gulwani;A. Meliou

文献摘要

被引文献

相似文献

数据驱动系统的推论的可靠性取决于数据的持续符合系统的初始设置和假设。当服务数据(我们要应用推论)偏离初始训练数据的轮廓时,推理的结果变得不可靠。我们介绍了符合性约束,这是一种针对量化不合格程度的新数据分析的原始限制,如果推断该元组是不可信的,则可以有效地表征。一致性约束是涉及数据集的数值属性的某些算术表达式(称为投影)的约束,现有的数据分析原始依据(例如功能依赖性和拒绝约束)无法建模。我们的关键发现是,在数据集构造有效符合约束的数据集上会产生较低的差异。该原理产生了一个令人惊讶的结果,即主成分分析的低变化成分通常被丢弃以减少维度,比高变化的组件产生更强的符合性约束。基于此结果,我们提供了高度可扩展和有效的技术 - 数据大小和属性数量的立方线 - 以发现数据集的一致性约束。为了衡量元组对数据集的不合格程度,我们提出了一种定量语义,该语义捕获了元组违反该数据集的一致性约束。我们证明了对两个应用程序的一致性约束的价值:值得信赖的机器学习和数据漂移。我们从经验上表明,一致性约束提供了(1)可靠地检测到不应信任机器学习模型的推断的元素,并且(2)比目前的状态更准确地量化数据漂移。
The reliability of inferences made by data-driven systems hinges on the data's continued conformance to the systems' initial settings and assumptions. When serving data (on which we want to apply inference) deviates from the profile of the initial training data, the outcome of inference becomes unreliable. We introduce conformance constraints, a new data profiling primitive tailored towards quantifying the degree of non-conformance, which can effectively characterize if inference over that tuple is untrustworthy. Conformance constraints are constraints over certain arithmetic expressions (called projections) involving the numerical attributes of a dataset, which existing data profiling primitives such as functional dependencies and denial constraints cannot model. Our key finding is that projections that incur low variance on a dataset construct effective conformance constraints. This principle yields the surprising result that low-variance components of a principal component analysis, which are usually discarded for dimensionality reduction, generate stronger conformance constraints than the high-variance components. Based on this result, we provide a highly scalable and efficient technique--linear in data size and cubic in the number of attributes--for discovering conformance constraints for a dataset. To measure the degree of a tuple's non-conformance with respect to a dataset, we propose a quantitative semantics that captures how much a tuple violates the conformance constraints of that dataset. We demonstrate the value of conformance constraints on two applications: trusted machine learning and data drift. We empirically show that conformance constraints offer mechanisms to (1) reliably detect tuples on which the inference of a machine-learned model should not be trusted, and (2) quantify data drift more accurately than the state of the art.