课题基金 / 基金详情

CAREER: Hashing and Sketching Algorithms for Resource-Frugal Machine Learning

CAREER: Hashing and Sketching Algorithms for Resource-Frugal Machine Learning
职业:用于资源节约型机器学习的哈希和草图算法
批准号:
1652131
负责人:
Anshumali Shrivastava
金额:
$49.91万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-05-01 至 2023-04-30

项目摘要

项目成果

Anshumali Shrivastava的其他基金

相似基金

相关文献

中文摘要
翻译
现代应用程序不断处理TB级的数据集,预计很快就会达到PB级。 当前数据集的大小和维度使得机器学习(ML)模型变得非常庞大和复杂,这增加了现有的问题。学习和推理的经典方法未能解决计算资源、存储限制、网络通信约束、能源效率、实时延迟本项目主要研究(指数)资源节约和可扩展的机器学习算法,非常适合当前的大数据约束。该项目利用概率哈希技术来推进状态-最先进的机器学习算法重点是重新设计现有的机器学习管道,使它们能够适应哈希加速。除了成本呈指数级下降外,所设计的算法还可以大规模并行化。三个主要目标是:1)通过哈希实现计算效率高的深度学习和基于内核的学习,2)(指数)压缩机器学习模型的草图算法,3)提高哈希函数的效率。该项目利用了几个最新的想法,包括非对称哈希,基于哈希的内核,加密哈希方案,次线性自适应采样和自适应草图,将学习算法推向极端规模。 通过在概率哈希和机器学习之间建立一个独特的桥梁,该项目进一步增强了当前对计算、空间和准确性权衡的理解。
英文摘要
Modern applications are constantly dealing with datasets at terabyte scale, and the anticipation is that very soon it will reach petabyte levels. The size and dimensionality of current datasets have made machine learning (ML) models significantly large and complex, which adds to the existing problems. Classical approaches to learning and inference fail to address new concerns of computational resources, storage limitations, network communication constraints, energy efficiency, real-time latency, etc. This project focuses on basic design and implementation of (exponentially) resource-frugal and scalable machine learning algorithms which are ideally suited for current big-data constraints.This project leverages probabilistic hashing techniques for advancing the state-of-the-art machine learning algorithms. The focus is on redesigning existing machine learning pipelines to make them amenable to the hashing speedup. Apart from being exponentially cheap, the designed algorithms are also massively parallelizable. The three primary objectives are: 1) Computationally Efficient Deep-Learning and Kernel-Based Learning via Hashing, 2) Sketching Algorithms for (Exponentially) Compressing Machine Learning Models, and 3) Improving Efficiency of Hash Functions. This project capitalizes on several recent ideas, including asymmetric hashing, hash-based kernels, densified hashing schemes, sub-linear adaptive sampling, and adaptive sketching, to push learning algorithms to the extreme-scale. By creating a unique bridge between probabilistic hashing and machine learning, this project further enhances the current understanding of tradeoffs involving computations, space, and accuracy.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3178876.3186056
发表时间: 2018-04
期刊: Proceedings of the 2018 World Wide Web Conference
影响因子: --
作者: [Chen Luo;Anshumali Shrivastava]
通讯作者: Chen Luo;Anshumali Shrivastava
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者: []
通讯作者:
DOI: 10.1145/3097983.3098035
发表时间: 2016-02
期刊: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子: --
作者: [Ryan Spring;Anshumali Shrivastava]
通讯作者: Ryan Spring;Anshumali Shrivastava
DOI: 10.1145/3183713.3196925
发表时间: 2018-05
期刊: Proceedings of the 2018 International Conference on Management of Data
影响因子: --
作者: [Yiqiu Wang;Anshumali Shrivastava;Jonathan Wang;Junghee Ryu]
通讯作者: Yiqiu Wang;Anshumali Shrivastava;Jonathan Wang;Junghee Ryu
共 7 条
    BIGDATA: F: Collaborative Research: Theory and Practice of Randomized Algorithms for Ultra-Large-Scale Signal Processing
    • 批准号:
      1838177
    • 项目类别:
      Standard Grant
    • 资助金额:
      $80.0万
    • 财政年份:
      2018
    • 负责人:
      Anshumali Shrivastava
    • 依托单位:
    海外基金