Approximate summaries for why and why-not provenance

Approximate summaries for why and why-not provenance
复制标题

DOI:
10.14778/3380750.3380760
复制
发表时间:
2020-01
影响因子:
2.5
通讯作者:
Seok-Gyun Lee;Bertram Ludäscher;Boris Glavic
Seok-Gyun Lee;Bertram Ludäscher;Boris Glavic
中科院分区:
计算机科学2区
文献类型:
--
作者:
Seok-Gyun Lee;Bertram Ludäscher;Boris Glavic

文献摘要

被引文献

相似文献

近年来,人们对“为什么”和“为什么”的来源进行了广泛的研究。然而,为什么不是来源,以及——在较小程度上——为什么来源可能非常大,从而导致严重的可伸缩性和可用性挑战。我们引入了一种新颖的近似概括技术来解决这些挑战。我们的方法使用模式来简洁地编码为什么和为什么不是来源。我们开发了有效计算来源摘要的技术,以平衡信息性、简洁性和完整性。为了实现可扩展性,我们将采样技术集成到来源捕获和汇总中。我们的方法是第一个既扩展到大型数据集,又生成全面而有意义的摘要的方法。
Why and why-not provenance have been studied extensively in recent years. However, why-not provenance and --- to a lesser degree --- why provenance can be very large, resulting in severe scalability and usability challenges. We introduce a novel approximate summarization technique for provenance to address these challenges. Our approach uses patterns to encode why and why-not provenance concisely. We develop techniques for efficiently computing provenance summaries that balance informativeness, conciseness, and completeness. To achieve scalability, we integrate sampling techniques into provenance capture and summarization. Our approach is the first to both scale to large datasets and generate comprehensive and meaningful summaries.