Selection Bias Tracking and Detailed Subset Comparison for High-Dimensional Data

Selection Bias Tracking and Detailed Subset Comparison for High-Dimensional Data
复制标题

DOI:
10.1109/tvcg.2019.2934209
复制
发表时间:
2019-06
影响因子:
5.2
通讯作者:
D. Borland;Wenyuan Wang;Jonathan Zhang;Joshua Shrestha;D. Gotz
D. Borland;Wenyuan Wang;Jonathan Zhang;Joshua Shrestha;D. Gotz
中科院分区:
计算机科学1区
文献类型:
--
作者:
D. Borland;Wenyuan Wang;Jonathan Zhang;Joshua Shrestha;D. Gotz

文献摘要

相似文献

大型、复杂数据集的收集在各个领域已经变得很常见。可视化分析工具在探索和回答有关这些大型数据集的复杂问题方面日益发挥着关键作用。然而,许多可视化并不是为了同时可视化复杂数据集中存在的大量维度(例如电子健康记录系统中的数万个不同代码)而设计的。这一事实,再加上许多可视化分析系统能够根据一小部分可视化维度对个体进行快速、临时的特定组或群组的能力,导致有可能引入选择偏差——当用户根据一组指定的维度创建群组时,也可能会引入许多其他看不见的维度之间的差异。这些意想不到的副作用可能会导致队列不再代表打算研究的更大人群,这可能会对后续分析的有效性产生负面影响。我们提出了选择偏差跟踪和可视化技术,可以将其合并到高维探索性视觉分析系统中,重点关注具有现有数据层次结构的医疗数据。这些技术包括:(1)基于树的队列起源和可视化,包括与所有其他队列进行比较的用户指定的基线队列,以及队列“漂移”的视觉编码,这表明可能发生选择偏差的位置,以及(2)一组可视化,包括基于新颖的冰柱图的可视化,以详细比较基线和用户指定的焦点队列之间的每维度差异。这些技术被集成到医学时间事件序列可视化分析工具中。我们提供示例用例并报告领域专家用户访谈的结果。
The collection of large, complex datasets has become common across a wide variety of domains. Visual analytics tools increasingly play a key role in exploring and answering complex questions about these large datasets. However, many visualizations are not designed to concurrently visualize the large number of dimensions present in complex datasets (e.g. tens of thousands of distinct codes in an electronic health record system). This fact, combined with the ability of many visual analytics systems to enable rapid, ad-hoc specification of groups, or cohorts, of individuals based on a small subset of visualized dimensions, leads to the possibility of introducing selection bias–when the user creates a cohort based on a specified set of dimensions, differences across many other unseen dimensions may also be introduced. These unintended side effects may result in the cohort no longer being representative of the larger population intended to be studied, which can negatively affect the validity of subsequent analyses. We present techniques for selection bias tracking and visualization that can be incorporated into high-dimensional exploratory visual analytics systems, with a focus on medical data with existing data hierarchies. These techniques include: (1) tree-based cohort provenance and visualization, including a user-specified baseline cohort that all other cohorts are compared against, and visual encoding of cohort “drift”, which indicates where selection bias may have occurred, and (2) a set of visualizations, including a novel icicle-plot based visualization, to compare in detail the per-dimension differences between the baseline and a user-specified focus cohort. These techniques are integrated into a medical temporal event sequence visual analytics tool. We present example use cases and report findings from domain expert user interviews.