Reptile: Aggregation-level Explanations for Hierarchical Data

Reptile: Aggregation-level Explanations for Hierarchical Data
复制标题

Reptile:分层数据的聚合级解释

DOI:
10.1145/3514221.3517854
复制
发表时间:
2022
期刊:
SIGMOD
影响因子:
--
通讯作者:
Wu, Eugene
Wu, Eugene
中科院分区:
--
文献类型:
--
作者:
Huang, Zezhou;Wu, Eugene

文献摘要

参考文献

相似文献

用户通常可以从概览级统计数据中看到一些结果看起来“不正常”,但很少能够描述错误的类型。Reptile是一个迭代的人在环解释和清理系统,用于分层数据中的错误。用户指定一个异常的分布式聚合结果(一个投诉),Reptile建议向下钻取操作来帮助用户“放大”潜在的错误。与先前干预原始记录的解释系统不同,Reptile通过学习一个组的预期统计数据进行干预,并根据干预修复投诉的程度对向下钻取的子组进行排名。这种组级公式支持广泛的错误类型(缺失、重复、值错误),并独特地利用了用户投诉的分布特性。此外,基于学习的干预允许用户提供Reptile学习的领域专业知识。在每次向下钻取迭代中,Reptile必须训练大量的预测模型。因此,我们扩展了因式学习从计数连接查询到聚合连接查询,并开发了一套优化,利用数据的层次结构。与基于Lapack的实现相比,这些优化将运行时间减少了>6倍。当应用于真实世界的Covid-19和非洲农民调查数据时,Reptile正确识别了21/30(使用现有解释方法为2)和20/22的错误。REPTILE已在埃塞俄比亚和赞比亚部署,并用于清理全国农民调查数据;干净的数据已用于设计国家干旱保险政策。
Users often can see from overview-level statistics that some results look "off", but are rarely able to characterize even the type of error. Reptile is an iterative human-in-the-loop explanation and cleaning system for errors in hierarchical data. Users specify an anomalous distributive aggregation result (a complaint), and Reptile recommends drill-down operations to help the user "zoom-in" on the underlying errors. Unlike prior explanation systems that intervene on raw records, Reptile intervenes by learning a group's expected statistics, and ranks drill-down sub-groups by how much the intervention fixes the complaint. This group-level formulation supports a wide range of error types (missing, duplicates, value errors) and uniquely leverages the distributive properties of the user complaint. Further, the learning-based intervention lets users provide domain expertise that Reptile learns from.In each drill-down iteration, Reptile must train a large number of predictive models. We thus extend factorized learning from count-join queries to aggregation-join queries, and develop a suite of optimizations that leverage the data's hierarchical structure. These optimizations reduce runtimes by >6× compared to a Lapack-based implementation. When applied to real-world Covid-19 and African farmer survey data, Reptile correctly identifies 21/30 (vs 2 using existing explanation approaches) and 20/22 errors. Reptile has been deployed in Ethiopia and Zambia, and used to clean nation-wide farmer survey data; the clean data has been used to design national drought insurance policies.
超越出处:用基于模式的平衡解释查询答案
DOI: 10.1145/3299869.3300066
发表时间: 2019
期刊: SIGMOD
影响因子: --
作者:
Miao, Zhengjie;Zeng, Qitian;Glavic, Boris;Roy, Sudeepa
通讯作者: Roy, Sudeepa
学校效能研究中的统计建模问题
DOI: --
发表时间: 1986
期刊:
影响因子: --
作者:
M. Aitkin;N. Longford
通讯作者: N. Longford
智能钻取:新的数据探索运算符
DOI: --
发表时间: 2014
影响因子: 2.5
作者:
Manas R. Joglekar;H. Garcia;Aditya G. Parameswaran
通讯作者: Aditya G. Parameswaran
人口多样性和不适应效应的动态多层次模型。
DOI: --
发表时间: 2005
影响因子: 9.9
作者:
J. Sacco;N. Schmitt
通讯作者: N. Schmitt
外在动机、PSM 和劳动力市场特征:26 个国家公共部门就业偏好的多层次模型
DOI: --
发表时间: 2015
期刊:
影响因子: --
作者:
S. Van de Walle;B. Steijn;S. Jilke
通讯作者: S. Jilke