On Detecting Cherry-picked Generalizations

On Detecting Cherry-picked Generalizations
复制标题

关于检测精挑细选的概括

DOI:
10.14778/3485450.3485457
复制
发表时间:
2021
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Tova Milo
Tova Milo
中科院分区:
--
文献类型:
--
作者:
Yin Lin;Brit Youngmann;Y. Moskovitch;H. V. Jagadish;Tova Milo

文献摘要

被引文献

相似文献

从详细的数据到更广泛的上下文中的语句的泛化对于用户理解大型数据集通常至关重要。相应地,结构不佳的概括可能传递误导性信息,即使这些陈述在技术上得到了数据的支持。例如,精挑细选的聚合水平可能会掩盖反对泛化的实质性子群。我们提出了一个框架,用于通过细化聚合查询来检测和解释精挑细选的泛化。我们提出了一种评分方法来表明推广的适当性。我们设计了高效的分数计算算法。为了更好地理解结果分数,我们还制定了实际的解释任务,以揭示重要的反例,并提供更好的替代陈述。我们使用真实世界的数据集和实例进行了实验,以显示我们所提出的评估度量的有效性和我们的算法框架的效率。
Generalizing from detailed data to statements in a broader context is often critical for users to make sense of large data sets. Correspondingly, poorly constructed generalizations might convey misleading information even if the statements are technically supported by the data. For example, a cherry-picked level of aggregation could obscure substantial sub-groups that oppose the generalization. We present a framework for detecting and explaining cherry-picked generalizations by refining aggregate queries. We present a scoring method to indicate the appropriateness of the generalizations. We design efficient algorithms for score computation. For providing a better understanding of the resulting score, we also formulate practical explanation tasks to disclose significant counterexamples and provide better alternatives to the statement. We conduct experiments using real-world data sets and examples to show the effectiveness of our proposed evaluation metric and the efficiency of our algorithmic framework.