Exploring the Effects of Aggregation Choices on Untrained Visualization Users' Generalizations From Data

Exploring the Effects of Aggregation Choices on Untrained Visualization Users' Generalizations From Data
复制标题

DOI:
10.1111/cgf.13902
复制
发表时间:
2020-02
影响因子:
2.5
通讯作者:
F. Nguyen;X. Qiao;Jeffrey Heer;J. Hullman
F. Nguyen;X. Qiao;Jeffrey Heer;J. Hullman
中科院分区:
计算机科学4区
文献类型:
--
作者:
F. Nguyen;X. Qiao;Jeffrey Heer;J. Hullman

文献摘要

相似文献

可视化系统设计者必须决定是否以及如何默认聚合数据。将分布信息聚合在单个摘要标记(例如平均值或总和)中可以简化解释,但可能会导致未经训练的用户忽略分布特征。我们问,未经训练的可视化用户得出的结论如何受到聚合策略的影响?我们提出了两个对照实验,比较未经训练的用户通过可视化对 1000 条记录或 50 条记录样本进行的可视化所做的概括,这些样本具有单个平均摘要标记、每个观察一个标记的分类视图或在分类视图之上覆盖平均摘要标记的视图。虽然我们观察到聚合策略对任一样本量的泛化准确性都没有可靠的影响,但纯粹分类视图的用户对其泛化的平均信心略低于视图显示单一平均摘要标记的用户,并且不太可能对存在或不存在的影响进行二分思考。比较 1000 条记录和 50 条记录数据集的结果,我们发现看到分类数据的观众相对于只看到平均摘要分数的观众,产生的概括数量和报告的概括信心有相当大的下降。
Visualization system designers must decide whether and how to aggregate data by default. Aggregating distributional information in a single summary mark like a mean or sum simplifies interpretation, but may lead untrained users to overlook distributional features. We ask, How are the conclusions drawn by untrained visualization users affected by aggregation strategy? We present two controlled experiments comparing generalizations of a population that untrained users made from visualizations that summarized either a 1000 record or 50 record sample with either single mean summary mark, a disaggregated view with one mark per observation or a view overlaying a mean summary mark atop a disaggregated view. While we observe no reliable effect of aggregation strategy on generalization accuracy at either sample size, users of purely disaggregated views were slightly less confident in their generalizations on average than users whose views show a single mean summary mark, and less likely to engage in dichotomous thinking about effects as either present or absent. Comparing results from 1000 record to 50 record data set, we see a considerably larger decrease in the number of generalizations produced and reported confidence in generalizations among viewers who saw disaggregated data relative to those who saw only mean summary marks.