To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested Data

To Not Miss the Forest for the Trees - A Holistic Approach for Explaining Missing Answers over Nested Data
复制标题

只见树木,不见森林——解释嵌套数据缺失答案的整体方法

DOI:
10.1145/3448016.3457249
复制
发表时间:
2021
期刊:
Proceedings of the 46th International Conference on Management of Data
影响因子:
--
通讯作者:
Glavic, Boris
Glavic, Boris
中科院分区:
--
文献类型:
--
作者:
Diestelkämper, Ralf;Lee, Seokki;Herschel, Melanie;Glavic, Boris

文献摘要

参考文献

被引文献

相似文献

对缺失答案的基于查询的解释确定查询的哪些运算符负责未能返回感兴趣的缺失答案。这种类型的解释已被证明是有用的,例如,调试复杂的分析查询。这种查询在Apache Spark等大数据系统中很常见。我们提出了一种新的方法来产生基于查询的解释。它是第一个支持嵌套数据并考虑修改数据模式和结构的运算符(例如,嵌套、投影)作为缺失答案的潜在原因。为了有效地计算解释,我们提出了一种启发式算法,采用两种新的技术:(i)推理多个模式的替代查询和(ii)重新验证在每一步是否中间结果可以有助于失踪的答案。使用Spark上的实现,我们证明了我们的方法是第一个扩展到大型数据集的方法,同时经常找到现有技术无法识别的解释。
Query-based explanations for missing answers identify which operators of a query are responsible for the failure to return a missing answer of interest. This type of explanations has proven useful, e.g., to debug complex analytical queries. Such queries are frequent in big data systems such as Apache Spark. We present a novel approach to produce query-based explanations. It is the first to support nested data and to consider operators that modify the schema and structure of the data (e.g., nesting, projections) as potential causes of missing answers. To efficiently compute explanations, we propose a heuristic algorithm that applies two novel techniques: (i) reasoning about multiple schema alternatives for a query and (ii) re-validating at each step whether an intermediate result can contribute to the missing answer. Using an implementation on Spark, we demonstrate that our approach is the first to scale to large datasets while often finding explanations that existing techniques fail to identify.
根据来源示例对联合查询进行逆向工程
DOI: --
发表时间: 2019
期刊: International Conference on Extending Database Technology
影响因子: --
作者:
Daniel Deutch;Amir Gilad
通讯作者: Amir Gilad
DOI: 10.1109/bigdata.2017.8258260
发表时间: 2017-12
期刊: 2017 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Pouria Pirzadeh;M. Carey;T. Westmann
通讯作者: Pouria Pirzadeh;M. Carey;T. Westmann
自适应模式数据库
DOI: --
发表时间: 2017
期刊: Online Proceedings
影响因子: --
作者:
Spoth, W.;Arab, B. S.;Chan, E. S.;Gawlick, D.;Ghoneimy, A.;Glavic, B.;Hammerschmidt, B.;Kennedy, O.;Lee, S.;Liu, Z. H.
通讯作者: Liu, Z. H.
用自然语言解释缺失的查询结果
DOI: --
发表时间: 2020
期刊: International Conference on Extending Database Technology
影响因子: --
作者:
Daniel Deutch;Nave Frost;Amir Gilad;Tomer Haimovich
通讯作者: Tomer Haimovich
嵌套数据基于查询的原因解释
DOI: --
发表时间: 2019
期刊: TaPP
影响因子: --
作者:
Diestelkämper, Ralf;Glavic, Boris;Herschel, Melanie;Lee, Seokki
通讯作者: Lee, Seokki