课题基金 / 基金详情

III: Small: An end-to-end pipeline for interactive visual analysis of big data

III: Small: An end-to-end pipeline for interactive visual analysis of big data
III:小型:用于大数据交互式可视化分析的端到端管道
批准号:
1815238
负责人:
Carlos Scheidegger
金额:
$48.58万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2021-08-31

项目摘要

项目成果

Carlos Scheidegger的其他基金

相似基金

相关文献

中文摘要
翻译
计算机科学家、统计学家和数据科学家使用复杂的分析技术从大量数据源中提取见解,从望远镜收集的天文数据到超级计算机上运行的气候模拟,再到在线社交网络中的用户活动。与此同时,他们希望利用交互式可视化,这样他们就可以通过图形和可视化界面来理解和探索他们的数据。这些可视化分析系统更直观,功能更强大,使分析师能够更自信地做出更好的决策。目前,这些可视化系统在大规模环境中的广泛适用性不够快。在这个项目中,开发了新的技术来加快数据分析中使用的方法,以便利益相关者将他们需要的复杂分析与他们喜欢使用的交互式可视化系统相结合。该项目有可能改变当前交互式和探索性数据分析基础设施和系统的设计方式。 将广泛传播直接与科学家和其他数据分析员使用的库和编程语言相结合的开放源码软件。此外,这里开发的概念和技术将在课堂上使用,以培训未来的研究人员和计算机科学家。目前,在大规模数据分析中应用交互式可视化系统的主要障碍是:许多技术需要在数据集上重复循环(或扫描),以收集适当的聚合信息。最近开发的分层时空数据立方体数据结构取代了许多扫描,但只适用于基本的条形图,直方图和热图,因为它们只加速了数据库管理系统可用的少量查询。相比之下,该项目旨在开发新型数据结构,支持更广泛的探索性数据分析和可视化管道,例如k均值、逻辑回归、最小二乘优化、降维等,并将这些数据结构直接连接到使用这些方法的可视化库所进行的API和调用。 互动和探索性数据分析的拟议基础设施的性能将根据专门设计的基准进行评估,以比较现有的和新的互动数据立方体系统。这些基准将使分散在一个有些支离破碎的研究领域的知识能够得到综合。这些基准将反过来指导对这些数据结构的改进开发的评估,旨在降低30%至80%的存储成本,并可能在预处理时间方面获得类似的收益,这直接转化为更好的交互式可视化功能。将开发并广泛传播用于将这些数据结构集成到R和Python等现代数据科学环境中的API,以提高该项目的影响力。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Computer scientists, statisticians, and data scientists use sophisticated analysis techniques to extract insights from their massive sources of data, from astronomic data gathered from telescopes to climate simulations run on supercomputers to user activity in online social networks. At the same time, they would like to make use of interactive visualization, so they can understand and explore their data by means of graphics and visual interfaces. These visual analytics systems are more intuitive and more powerful, and allow analysts to make better decisions more confidently. Currently, these visualization systems are not fast enough for broad applicability in large-scale settings. In this project, novel techniques are developed to speed up the methods used in data analyses in order for stakeholders to combine the sophisticated analyses they need with the interactive visualization systems they prefer to use. This project has the potential to transform how current infrastructure and systems for interactive and exploratory data analysis are designed. Open-source software that integrates directly with the libraries and programming languages used by scientists and other data analysts will be broadly disseminated. In addition, the concepts and technologies developed here will be used in classrooms to train future generations of researchers and computer scientists.There currently is a major obstacle for the application of interactive visualization systems in large-scale data analysis: many techniques require repeated loops (or scans) over the dataset in order to collect the appropriate aggregation information. The recently developed hierarchical, spatiotemporal data cube data structures replace many of scans, but are only suitable for basic bar charts, histograms, and heatmaps, since they accelerate only a small number of queries available to database management systems. In contrast, this project aims at developing novel data structures that support a broader swath of the exploratory data analysis and visualization pipeline, such as k-means, logistic regression, least-squares optimization, dimensionality reduction, etc., and connect these data structures directly to the APIs and calls made by visualization libraries that use these methods. The performance of proposed infrastructure for interactive and exploratory data analysis will be evaluated on specifically designed benchmarks to compare existing and novel interactive data cube systems. The benchmarks will enable synthesis of knowledge that is spread across a somewhat fractured research area. The benchmarks will, in turn, guide the evaluation of the development of improvement for these data structures, aiming at a decrease between 30% to 80% in storage costs, and likely comparable gains in preprocessing time, that translate directly into better interactive visualization capabilities. APIs for integrating these data structures in modern data science environments such as R and Python will be developed and widely disseminated in order to increase the impact of this project.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Disentangling Influence: Using disentangled representations to audit model predictions
解缠结影响:使用解缠结表示来审核模型预测
DOI: --
发表时间: 2019
期刊: Proceedings of Neural Information Processing Systems (NeurIPS
影响因子: --
作者: [Marx, Charles, Phillips, Richard, Friedler, Sorelle A., Scheidegger, Carlos, Venkatasubramanian, Suresh]
通讯作者: Venkatasubramanian, Suresh
DOI: 10.1109/tvcg.2020.3028891
发表时间: 2020-10
期刊: IEEE Transactions on Visualization and Computer Graphics
影响因子: 5.2
作者: [L. Battle;C. Scheidegger]
通讯作者: L. Battle;C. Scheidegger
III: Medium: Collaborative Research: Evaluating and Maximizing Fairness in Information Flow on Networks
  • 批准号:
    1955162
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $25.04万
  • 财政年份:
    2020
  • 负责人:
    Carlos Scheidegger
  • 依托单位:
III: Medium: Collaborative Research: Topological Data Analysis for Large Network Visualization
  • 批准号:
    1513651
  • 项目类别:
    Standard Grant
  • 资助金额:
    $26.89万
  • 财政年份:
    2015
  • 负责人:
    Carlos Scheidegger
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: