课题基金 / 基金详情

TensorLABE - Robust Characterization of Data Tensors and Synthetic Data Generation

TensorLABE - Robust Characterization of Data Tensors and Synthetic Data Generation
TensorLABE - 数据张量的稳健表征和合成数据生成
批准号:
2223932
负责人:
Tim Andersen
金额:
$15.65万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-09-01 至 2024-08-31

项目摘要

项目成果

Tim Andersen的其他基金

相似基金

相关文献

中文摘要
翻译
现代计算方法,如基于机器学习(ML)的方法,在效率和性能方面取得了令人印象深刻的进展,但越来越依赖于海量数据。这些数据驱动的方法从人类工程源代码的经典技术过渡到在数据集上训练的算法,以产生所需的解决方案,将数据放在驾驶座上。新的硬件和软件系统专门为支持与这些算法相关的复杂数据驱动计算以及伴随它们的海量数据而设计,从而推动和加速了这些数据驱动技术的扩散。但是,尽管这些新的硬件和软件系统获得了令人印象深刻的性能提升,但对其设计的数据组件的理解一直停滞不前,而是支持基于软件和硬件的解决方案中以性能为驱动的进步。缺乏对数据的理解导致了许多不良结果,例如数据驱动的解决方案中不想要的偏见、无法确定数据集对于提前解决给定问题的实际适宜性、无法确定数据集是否被操纵或损坏、以及无法产生可用于训练和测试这些软件和硬件系统的性能的准确的合成数据。该项目旨在为基于张量的大规模数据集的表征提供一个稳健的框架,以提高对数据本身的理解,并使合成数据能够更准确地复制真实世界的数据,用于系统设计测试和验证。具体地说,该项目建议通过结合用于统计、结构和执行数据分析的各种张量方法来提升多线性代数、大规模数据分析、机器学习和人工智能领域的知识,以实现更稳健的数据表征。一套更全面的数据特征将能够更好地评估数据的偏差和评估数据集对特定任务的适宜性。它还将允许对数据集进行比较,以了解它们的差异,并评估数据是否存在腐败或操纵。通过将项目中开发的数据表征方法结合到生成比传统方法更真实的合成数据中,将建立概念证明。该方法将通过测试合成数据更准确地表征软件/硬件系统性能的能力进行验证。该奖项反映了NSF的法定使命,并已通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern computational methods, such as Machine Learning (ML) based approaches, have produced impressive gains in efficiency and performance but are increasingly dependent on massive amounts of data. These data-driven approaches transition from the classical techniques of human-engineered source code to algorithms trained on a dataset to produce the desired solution, placing the data in the driver's seat. The proliferation of these data-driven technologies is being enabled and hastened by new hardware and software systems specifically designed to support the complex data-driven computation associated with these algorithms and the massive volumes of data accompanying them. But despite the impressive performance gains of these new hardware and software systems, understanding their design's data component has languished in favor of performance-driven advances in software and hardware-based solutions. The lack of data understanding has led to a number of undesirable outcomes such as unwanted bias in the data-driven solution, an inability to determine the actual suitability of a data set to solving a given problem ahead of time, an inability to determine if a data set has been manipulated or corrupted, and an inability to produce accurate synthetic data that can be used to train and test the performance of these software and hardware systems. This project aims to provide a robust framework for the characterization of large-scale tensor-based datasets to improve understanding of the data itself and enable the production of synthetic data that more accurately replicates real-world data for use in system design testing and validation.Specifically, this project proposes to advance knowledge in the fields of multilinear algebra, large-scale data analytics, machine learning, and artificial intelligence by incorporating a variety of tensor methods for statistical, structural, and performative data analyses to achieve more robust data characterization. A more holistic set of data characterizations will enable better assessment of data for bias and evaluation of datasets for suitability for a particular task. It will also allow the comparison of datasets to understand their differences and assess data for corruption or manipulation. A proof of concept will be established by incorporating the data characterization methods developed in the project into generating synthetic data with higher degrees of realism than conventional methods. The approach will be validated by testing the ability of the synthetic data to characterize software/hardware system performance more accurately.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SemiSynBio: Nucleic Acid Memory
  • 批准号:
    1807809
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $112.5万
  • 财政年份:
    2018
  • 负责人:
    Tim Andersen
  • 依托单位:
EAGER: Tensor500: A Streaming Analytics High Performance Computing Benchmark
  • 批准号:
    1849463
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.99万
  • 财政年份:
    2018
  • 负责人:
    Tim Andersen
  • 依托单位:
EAGER: Stream500: A New Benchmark and Infrastructure for Streaming Analytics
  • 批准号:
    1641774
  • 项目类别:
    Standard Grant
  • 资助金额:
    $11.95万
  • 财政年份:
    2016
  • 负责人:
    Tim Andersen
  • 依托单位:
国内基金
海外基金
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
  • 批准号:
    70601028
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    7.0万元
  • 批准年份:
    2006
  • 负责人:
    王明征
  • 依托单位:
心理紧张和应力影响下Robust语音识别方法研究
  • 批准号:
    60085001
  • 项目类别:
    专项基金项目
  • 资助金额:
    14.0万元
  • 批准年份:
    2000
  • 负责人:
    韩纪庆
  • 依托单位:
ROBUST语音识别方法的研究
  • 批准号:
    69075008
  • 项目类别:
    面上项目
  • 资助金额:
    3.5万元
  • 批准年份:
    1990
  • 负责人:
    高雨青
  • 依托单位:
改进型ROBUST序贯检测技术
  • 批准号:
    68671030
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1986
  • 负责人:
    刘有恒
  • 依托单位: