Explaining Aggregates for Exploratory Analytics

Explaining Aggregates for Exploratory Analytics
复制标题

DOI:
10.1109/bigdata.2018.8621953
复制
发表时间:
2018-10
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Fotis Savva;C. Anagnostopoulos;P. Triantafillou
Fotis Savva;C. Anagnostopoulos;P. Triantafillou
中科院分区:
其他
文献类型:
--
作者:
Fotis Savva;C. Anagnostopoulos;P. Triantafillou

文献摘要

被引文献

相似文献

希望探索多变量数据空间的分析人员通常会提出涉及选择操作符的查询,即范围或半径查询,这些查询定义可能感兴趣的数据子空间,然后使用聚合函数,其结果决定了他们的探索分析兴趣。然而,这样的聚合查询(AQ)结果是简单的标量,因此,对于探索性分析来说,传递的关于查询子空间的信息有限。我们通过提供一种新的解释机制来解决这个缺点,帮助分析人员探索和理解数据子空间,这种机制被称为XAXA:解释探索性分析的聚合。XAXA的新AQ解释是用一个三重联合优化问题得到的函数来表示的。解释采用一组参数分段线性函数的形式,通过统计学习模型获得。所提出的解决方案的一个关键特征是,模型训练只通过在线监测问题及其答案来执行。在XAXA中,可以在不访问任何数据库(DB)的情况下计算未来问题的解释,并且可以用于进一步探索查询的数据子空间,而无需向DB发出任何查询。我们通过对真实世界和合成数据集以及查询工作负载的理论基础指标来评估XAXA的解释准确性和效率。
Analysts wishing to explore multivariate data spaces, typically pose queries involving selection operators, i.e., range or radius queries, which define data subspaces of possible interest and then use aggregation functions, the results of which determine their exploratory analytics interests. However, such aggregate query (AQ) results are simple scalars and as such, convey limited information about the queried subspaces for exploratory analysis. We address this shortcoming aiding analysts to explore and understand data subspaces by contributing a novel explanation mechanism coined XAXA: eXplaining Aggregates for eXploratory Analytics. XAXA’s novel AQ explanations are represented using functions obtained by a three-fold joint optimization problem. Explanations assume the form of a set of parametric piecewise-linear functions acquired through a statistical learning model. A key feature of the proposed solution is that model training is performed by only monitoring AQs and their answers on-line. In XAXA, explanations for future AQs can be computed without any database (DB) access and can be used to further explore the queried data subspaces, without issuing any more queries to the DB. We evaluate the explanation accuracy and efficiency of XAXA through theoretically grounded metrics over real-world and synthetic datasets and query workloads.