课题基金 / 基金详情

III: Medium: Bias Tracking and Reduction Methods for High-Dimensional Exploratory Visual Analysis and Selection

III: Medium: Bias Tracking and Reduction Methods for High-Dimensional Exploratory Visual Analysis and Selection
III:中:高维探索性视觉分析和选择的偏差跟踪和减少方法
批准号:
1704018
负责人:
David Gotz
金额:
$108.16万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2022-11-30

项目摘要

项目成果

David Gotz的其他基金

相似基金

相关文献

中文摘要
翻译
大型复杂数据集的探索性可视化和分析在各个领域越来越普遍。例如,在线公司跟踪用户以了解他们的产品,计算机安全日志捕获网络活动的详细踪迹,医疗保健系统捕获患者的详细纵向记录。在所有这些领域,正在创建大型而复杂的数据存储库,目的是支持数据驱动的、基于证据的决策。然而,今天的可视化工具——分析师工具箱的关键部分——在应用于高维数据集(即具有大量变量的数据集)时往往不堪重负。现实世界的数据集通常有数千个变量;这与大多数可视化所支持的维度数量少得多形成鲜明对比。这种维度上的差距使任何分析的有效性都有很大的偏差风险,可能导致严重的、隐藏的错误。本研究项目将开发一种新的高维探索性可视化方法,有助于在探索性高维数据可视化过程中发现和减少选择偏差和其他数据解释问题。该项目的成果,包括开源软件,将广泛适用于各个领域。此外,该项目将在健康结果研究环境中与用户一起进行评估。这为改善世界各地的卫生保健提供了巨大的潜力。该项目开发了一套用于探索性数据分析的上下文可视化方法,旨在支持从高维数据中发现更强大和可推广的见解。这些方法建立在这样一种认识的基础上,即使许多视觉方法有效的摘要本身也会模糊高维数据集的某些方面,而这些方面可能对准确解释用户的视觉发现至关重要。更具体地说,在可视化中主动处理的数据子集(包括维度和记录)——数据焦点——必须在许多维度和数据记录的上下文中进行解释,这些维度和数据记录在可视化中被省略或没有清楚地表示——数据上下文中。因此,本项目开发的方法旨在(1)明确建模和分析数据上下文,(2)传达数据焦点与上下文之间的关系,以便更好地告知用户诸如混杂变量和选择偏差等隐藏问题。该项目的主要技术贡献包括:(1)用于视觉验证的内联复制;(2)高维可视化的基线选择方法;(3)代表性可视化交互式再平衡。此外,开源软件将被开发,并与现实世界的数据和从业者进行评估。该研究项目的产品——包括新方法、软件产品和评估结果——将通过项目网站(https://vaclab.web.unc.edu/contextual-visualization/)发布。
英文摘要
Exploratory visualization and analysis of large and complex datasets is growing increasingly common across a range of domains. For example, online companies track users to learn about their products, computer security logs capture detailed traces of network activity, and health care systems capture detailed longitudinal records for their patients. In all of these fields, large and complex data repositories are being created with the goal supporting data-driven, evidence-based decision making. However, today's visualization tools -- a critical part of an analyst's toolbox -- are often overwhelmed when applied to high-dimensional datasets (i.e., datasets with large numbers of variables). Real-world datasets can often have many thousands of variables; a stark contrast to the much smaller number of dimensions supported by most visualizations. This gap in dimensionality puts the validity of any analysis at great risk of bias, potentially leading to serious, hidden errors. This research project will develop a new approach to high-dimensional exploratory visualization that will help detect and reduce selection bias and other problems with data interpretation during exploratory high-dimensional data visualization. The project's results, including open-source software, will be broadly applicable across domains. In addition, the project will be evaluated with users in a health outcomes research setting. This offers significant potential to improve health care around the world. This project develops a set of Contextual Visualization Methods for exploratory data analysis which are designed to support the discovery of more robust and generalizable insights from high-dimensional data. These methods are built upon a recognition that the very summarization that makes many visual methods effective also inherently obscures aspects of a high-dimensional dataset that may be critical to accurate interpretation of a user's visual findings. More specifically, the subset of data (comprising both dimensions and records) that is actively accounted for within a visualization -- the data focus -- must be interpreted within the context of the many dimensions and data records that have been omitted or are not clearly represented within a visualization--the data context. The methods that this project develops, therefore, are designed to (1) explicitly model and analyze the data context, and (2) convey the relationship between the data focus and the context in order to better inform users about hidden problems such as confounding variables and selection bias. The primary technical contributions of the project include: (1) inline replication for visual validation; (2) baselined selection methods for high-dimensional visualization; (3) interactive rebalancing for representative visualization. In addition, open-source software will be developed and evaluated with real-world data and practitioners. The products of this research project -- including new methods, software products, and evaluation results -- will be disseminated through a project website (https://vaclab.web.unc.edu/contextual-visualization/).
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tvcg.2019.2934209
发表时间: 2019-06
期刊: IEEE Transactions on Visualization and Computer Graphics
影响因子: 5.2
作者: [D. Borland;Wenyuan Wang;Jonathan Zhang;Joshua Shrestha;D. Gotz]
通讯作者: D. Borland;Wenyuan Wang;Jonathan Zhang;Joshua Shrestha;D. Gotz
Enabling Longitudinal Exploratory Analysis of Clinical COVID Data
实现临床 COVID 数据的纵向探索性分析
DOI: 10.1109/vahc53616.2021.00008
发表时间: 2021
期刊: Proceedings of Visual Analytics in Healthcare (VAHC
影响因子: --
作者: [Borland, David, Brain, Irena, Fecho, Karamarie, Pfaff, Emily, Xu, Hao, Champion, James, Bizon, Chris, Gotz, David]
通讯作者: Gotz, David
Adaptive Contextualization Methods for Combating Selection Bias during High-Dimensional Visualization
在高维可视化过程中对抗选择偏差的自适应情境化方法
DOI: 10.1145/3009973
发表时间: 2017
期刊: ACM Transactions on Interactive Intelligent Systems
影响因子: 3.4
作者: [Gotz, David, Sun, Shun, Cao, Nan, Kundu, Rita, Meyer, Anne-Marie]
通讯作者: Meyer, Anne-Marie
DOI: 10.1109/tvcg.2019.2934661
发表时间: 2019-06
期刊: IEEE Transactions on Visualization and Computer Graphics
影响因子: 5.2
作者: [D. Gotz;Jonathan Zhang;Wenyuan Wang;Joshua Shrestha;D. Borland]
通讯作者: D. Gotz;Jonathan Zhang;Wenyuan Wang;Joshua Shrestha;D. Borland
11
    III: Medium: Counterfactual-Based Supports For Visual Causal Inference
    NSF Student Travel Support for the 2019 IEEE Visualization Doctoral Colloquium (IEEE VIS DC)
    QuBBD: Collaborative Research: Interactive Ensemble clustering for mixed data with application to mood disorders
    海外基金