课题基金 / 基金详情

CAREER: Multi-Query Optimizations for Deep Learning Systems

CAREER: Multi-Query Optimizations for Deep Learning Systems
职业:深度学习系统的多查询优化
批准号:
1942724
负责人:
Arun Kumar
金额:
$55.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-07-01 至 2025-06-30

项目摘要

项目成果

Arun Kumar的其他基金

相似基金

相关文献

中文摘要
翻译
使用被称为“深度学习”的预测模型的大规模数据分析已经彻底改变了许多数字应用程序,为现代语音识别、语言翻译、网络搜索等提供了动力。深度学习的成功(主要是在资源丰富的技术公司)引起了人们对在领域科学、企业公司、医疗保健甚至数字人文学科中采用深度学习的浓厚兴趣。但广泛采用深度学习的一个主要瓶颈是训练深度学习模型的高资源成本,这需要一个计算成本很高的经验过程,需要大量的试验。这个缓慢的过程增加了资源成本,浪费了能源,并阻碍了用户的生产力。该项目通过设计新技术来解决这个问题,从而在深度学习系统上大大加快这一过程。它将降低资源成本和能源需求,反过来,有助于将深度学习大众化到更多的应用领域。它将导致一个新的开源系统与现有流行的深度学习工具集成,使其更便宜,更快,更容易采用大规模深度学习。该系统将被领域科学家使用,也将集成到工业产品中。这项研究将通过顶级会议的出版物传播,并纳入数据分析系统的新课程。该项目将支持研究生、本科生和高中生,包括LGBT+和女学生。该项目将提高可扩展深度学习模型选择的资源效率,这是一个经验过程,通常需要训练数十到数百个具有不同数据表示、神经架构和超参数值的模型配置。像TensorFlow和PyTorch这样的现有工具专注于一次训练一个模型的效率,这在模型选择过程中浪费了大量的资源。一些系统还牺牲了再现性,这是许多用户的一大障碍。这个项目通过提出一种新的数据库系统启发的深度学习视图来解决这些问题,将其执行重新想象为查询。针对小型集群设置,它将三种常见的深度学习模型选择任务的规范提升到声明性级别,并一次运行许多相关的模型配置。它提出了一套多查询优化和视图物化技术,以减少通信成本和/或避免计算冗余,同时不牺牲再现性或预测准确性。这些技术将随机梯度下降的数学特性和深度学习查询的计算特性与仔细的并行数据系统设计和实现相结合。项目网站:https://adalabucsd.github.io/cerebrosystem/This该奖项反映了美国国家科学基金会的法定使命,并通过基金会的知识价值和更广泛的影响审查标准进行评估,认为值得支持。
英文摘要
Large-scale data analytics using predictive models called "deep learning" has revolutionized many digital applications, powering modern speech recognition, language translation, Web search, and more. This success of deep learning, primarily at resource-rich technology companies, has led to high interest in adopting deep learning in domain sciences, enterprise companies, healthcare, and even digital humanities. But a major bottleneck to broader adoption is the high resource cost of training deep learning models, which requires a computationally expensive empirical process with a large number of trials. This slow process raises resource costs, wastes energy, and impedes user productivity. This project tackles this problem by devising new techniques to substantially speedup up this process on deep learning systems. It will reduce resource costs and energy needs, and in turn, help democratize deep learning to more application domains. It will lead to a new open source system integrated with existing popular deep learning tools to make it cheaper, faster, and easier to adopt large-scale deep learning. The system will be used by domain scientists and also integrated into industrial products. The research will be disseminated via publications at top conferences and incorporated into new courses on data analytics systems. This project will support graduate, undergraduate, and high school students, including LGBT+ and female students.This project will improve the resource efficiency of scalable deep learning model selection, an empirical process that typically requires training dozens to hundreds of model configurations with varying data representations, neural architectures, and hyper-parameter values. Existing tools like TensorFlow and PyTorch focus on the efficiency of training one model a time, which wastes resources at scale during model selection. Some systems also sacrifice reproducibility, a showstopper for many users. This project resolves these issues by presenting a fresh database systems-inspired view of deep learning that re-imagines its executions as queries. Targeting small cluster settings, it raises the specification of three common deep learning model selection tasks to a declarative level and runs many related model configurations in one go. It proposes a suite of multi-query optimization and view materialization techniques that reduce communication costs and/or avoid computational redundancy, while not sacrificing reproducibility or prediction accuracy. The techniques combine the mathematical properties of stochastic gradient descent and the computational properties of deep learning queries with careful parallel data system design and implementation. Project website: https://adalabucsd.github.io/cerebrosystem/This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Arun Kumar;Advitya Gemawat;Kabir;Nagrecha;Yuhao Zhang;Side Li]
通讯作者: Arun Kumar;Advitya Gemawat;Kabir;Nagrecha;Yuhao Zhang;Side Li
DOI: 10.14778/3476249.3476284
发表时间: 2021-07
期刊: Proc. VLDB Endow.
影响因子: --
作者: [Side Li;Arun Kumar]
通讯作者: Side Li;Arun Kumar
Cerebro: a data system for optimized deep learning model selection
Cerebro:用于优化深度学习模型选择的数据系统
DOI: 10.14778/3407790.3407816
发表时间: 2020
期刊: Proceedings of the VLDB Endowment
影响因子: 2.5
作者: [Nakandala, Supun, Zhang, Yuhao, Kumar, Arun]
通讯作者: Kumar, Arun
DOI: 10.1145/3514221.3517846
发表时间: 2022-06
期刊: Proceedings of the 2022 International Conference on Management of Data
影响因子: --
作者: [Supun Nakandala;Arun Kumar]
通讯作者: Supun Nakandala;Arun Kumar
7
    III: Small: Towards Speech-Driven Multimodal Querying
    • 批准号:
      1816701
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2018
    • 负责人:
      Arun Kumar
    • 依托单位:
    国内基金
    海外基金
    基于Multi-Pass Cell的高功率皮秒激光脉冲非线性压缩关键技术研究
    Multi-decadeurbansubsidencemonitoringwithmulti-temporaryPStechnique
    • 批准号:
      --
    • 项目类别:
      --
    • 资助金额:
      80万元
    • 批准年份:
      2022
    • 负责人:
      Timo Balz
    • 依托单位:
    High-precision force-reflected bilateral teleoperation of multi-DOF hydraulic robotic manipulators
    • 批准号:
      52111530069
    • 项目类别:
      国际(地区)合作与交流项目
    • 资助金额:
      10万元
    • 批准年份:
      2021
    • 负责人:
      徐兵
    • 依托单位:
    大地电磁强噪音压制的Multi-RRMC技术及其在青藏高原东南缘-印支块体地壳流追踪中的应用