课题基金 / 基金详情

CAREER: Multi-Query Optimizations for Deep Learning Systems

CAREER: Multi-Query Optimizations for Deep Learning Systems
职业:深度学习系统的多查询优化
批准号:
1942724
负责人:
Arun Kumar
金额:
$55.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-07-01 至 2025-06-30

项目摘要

项目成果

Arun Kumar的其他基金

相似基金

相关文献

中文摘要
翻译
使用被称为“深度学习”的预测模型的大规模数据分析已经彻底改变了许多数字应用,推动了现代语音识别、语言翻译、网络搜索等。深度学习的这种成功,主要是在资源丰富的技术公司,导致了在领域科学、企业公司、医疗保健甚至数字人文领域采用深度学习的浓厚兴趣。但是,更广泛采用的一个主要瓶颈是训练深度学习模型的高资源成本,这需要通过大量试验进行计算昂贵的经验过程。这一缓慢的过程增加了资源成本,浪费了能源,并阻碍了用户的生产力。这个项目通过设计新的技术来大幅加速深度学习系统上的这一过程,从而解决了这个问题。它将降低资源成本和能源需求,并反过来帮助将深度学习民主化到更多的应用领域。它将导致一个新的开源系统,与现有的流行深度学习工具相集成,使其更便宜、更快、更容易采用大规模深度学习。该系统将由领域科学家使用,并集成到工业产品中。这项研究将通过顶级会议的出版物传播,并纳入关于数据分析系统的新课程。该项目将支持研究生、本科生和高中生,包括LGBT和女性学生。该项目将提高可扩展深度学习模型选择的资源效率,这是一个经验过程,通常需要训练数十到数百个具有不同数据表示、神经结构和超参数值的模型配置。现有的工具如TensorFlow和PyTorch专注于一次训练一个模型的效率,这在模型选择过程中大规模地浪费资源。一些系统还牺牲了重复性,这对许多用户来说是一个令人惊叹的问题。该项目通过提供一种受数据库系统启发的深度学习的新视图来解决这些问题,该视图将其执行重新想象为查询。针对小型集群环境,将三种常见深度学习模型选择任务的规范提升到声明性层次,并一次性运行多个相关的模型配置。它提出了一套多查询优化和视图物化技术,在不牺牲重复性或预测精度的情况下,降低了通信成本和/或避免了计算冗余。该技术结合了随机梯度下降的数学特性和深度学习查询的计算特性,并结合了仔细的并行数据系统的设计和实现。项目网站:https://adalabucsd.github.io/cerebrosystem/This奖反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Large-scale data analytics using predictive models called "deep learning" has revolutionized many digital applications, powering modern speech recognition, language translation, Web search, and more. This success of deep learning, primarily at resource-rich technology companies, has led to high interest in adopting deep learning in domain sciences, enterprise companies, healthcare, and even digital humanities. But a major bottleneck to broader adoption is the high resource cost of training deep learning models, which requires a computationally expensive empirical process with a large number of trials. This slow process raises resource costs, wastes energy, and impedes user productivity. This project tackles this problem by devising new techniques to substantially speedup up this process on deep learning systems. It will reduce resource costs and energy needs, and in turn, help democratize deep learning to more application domains. It will lead to a new open source system integrated with existing popular deep learning tools to make it cheaper, faster, and easier to adopt large-scale deep learning. The system will be used by domain scientists and also integrated into industrial products. The research will be disseminated via publications at top conferences and incorporated into new courses on data analytics systems. This project will support graduate, undergraduate, and high school students, including LGBT+ and female students.This project will improve the resource efficiency of scalable deep learning model selection, an empirical process that typically requires training dozens to hundreds of model configurations with varying data representations, neural architectures, and hyper-parameter values. Existing tools like TensorFlow and PyTorch focus on the efficiency of training one model a time, which wastes resources at scale during model selection. Some systems also sacrifice reproducibility, a showstopper for many users. This project resolves these issues by presenting a fresh database systems-inspired view of deep learning that re-imagines its executions as queries. Targeting small cluster settings, it raises the specification of three common deep learning model selection tasks to a declarative level and runs many related model configurations in one go. It proposes a suite of multi-query optimization and view materialization techniques that reduce communication costs and/or avoid computational redundancy, while not sacrificing reproducibility or prediction accuracy. The techniques combine the mathematical properties of stochastic gradient descent and the computational properties of deep learning queries with careful parallel data system design and implementation. Project website: https://adalabucsd.github.io/cerebrosystem/This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者: [Arun Kumar;Advitya Gemawat;Kabir;Nagrecha;Yuhao Zhang;Side Li]
通讯作者: Arun Kumar;Advitya Gemawat;Kabir;Nagrecha;Yuhao Zhang;Side Li
DOI: 10.14778/3476249.3476284
发表时间: 2021-07
期刊: Proc. VLDB Endow.
影响因子: --
作者: [Side Li;Arun Kumar]
通讯作者: Side Li;Arun Kumar
Cerebro: a data system for optimized deep learning model selection
Cerebro:用于优化深度学习模型选择的数据系统
DOI: 10.14778/3407790.3407816
发表时间: 2020
期刊: Proceedings of the VLDB Endowment
影响因子: 2.5
作者: [Nakandala, Supun, Zhang, Yuhao, Kumar, Arun]
通讯作者: Kumar, Arun
DOI: 10.1145/3514221.3517846
发表时间: 2022-06
期刊: Proceedings of the 2022 International Conference on Management of Data
影响因子: --
作者: [Supun Nakandala;Arun Kumar]
通讯作者: Supun Nakandala;Arun Kumar
7
    III: Small: Towards Speech-Driven Multimodal Querying
    • 批准号:
      1816701
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2018
    • 负责人:
      Arun Kumar
    • 依托单位:
    国内基金
    海外基金
    基于Multi-Pass Cell的高功率皮秒激光脉冲非线性压缩关键技术研究
    Multi-decadeurbansubsidencemonitoringwithmulti-temporaryPStechnique
    • 批准号:
      --
    • 项目类别:
      --
    • 资助金额:
      80万元
    • 批准年份:
      2022
    • 负责人:
      Timo Balz
    • 依托单位:
    High-precision force-reflected bilateral teleoperation of multi-DOF hydraulic robotic manipulators
    • 批准号:
      52111530069
    • 项目类别:
      国际(地区)合作与交流项目
    • 资助金额:
      10万元
    • 批准年份:
      2021
    • 负责人:
      徐兵
    • 依托单位:
    大地电磁强噪音压制的Multi-RRMC技术及其在青藏高原东南缘-印支块体地壳流追踪中的应用