Computational Foundations of Machine Learning in the Era of Big Data
Computational Foundations of Machine Learning in the Era of Big Data
批准号:
RGPIN-2017-05032
负责人:
Yu, Yaoliang
金额:
$2.04万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
机器学习(ML)是一个开发能够通过学习和经验改进自身的软件的领域,它在很大程度上是由历史数据的可用性以及开发高效和可扩展的算法和支持理论的需要推动的。相反,ML在科学、工程和商业领域的成功,以及技术创新,导致了大数据收集领域前所未有的增长和热情,从而重新定义了计算效率并吸引了系统解决方案。例如,最近DeepMind的AlphaGo系统击败了顶级人类围棋选手,需要1900个CPU和280个GPU来进行计算。如何在这个庞大的分布式集群中平衡计算与通信,而不影响系统吞吐量或正确性?另一方面,一家开发移动应用程序的小型初创公司可能负担不起与谷歌相同的计算能力,因此经常不得不变成原始的解决方案。如何为ML构建一个算法框架,提供“旋钮”来调整计算负载,并在精度上显式地、可控地损失?因此,满足大数据时代如此多样化的计算需求对ML领域来说是一个巨大的挑战。*我们试图通过三个相辅相成的目标来解决ML和大数据中的这种计算挑战:(1)真正的问题是困难的,但也是结构化的。多年来,设计能够利用数据和模型中的某些结构的统计方法和计算算法的重要性已变得显而易见。在我们之前关于稀疏性和低秩性的工作的鼓励下,我们建议研究ML应用中常见的另外两种结构:单调性和多模态(张量格式),并开发受益于这些结构的高效算法。(2)数据总是带有噪声和随机波动,从而减少了用最大似然方法获得精确甚至高精度解的需要。如果处理得当,近似计算可以显著减少ML的计算时间。我们开始系统地研究ML中近似计算的权衡,从计算代价昂贵的程序降级到更简单、更廉价的程序,再到最优的光滑不可微函数,并将非凸性的度量附加到非凸函数上。(3)分布式计算已成为处理大数据集的标准。为了更好地平衡分布式ML系统中的通信和计算,我们提出了有界异步协议(BAP),并继续研究了典型ML迭代算法在BAP和可能不那么严格的凸或光滑假设下的加速比和收敛保证。我们的工作将进一步推进ML的计算理论和实践,所产生的算法和系统将是使用ML方法分析大数据集的基础。
英文摘要
Machine learning (ML), a field that develops software that can improve itself through learning and experience, has been largely driven by the availability of historical data, and by the need to develop efficient and scalable algorithms and supporting theories. Conversely, the success of ML in science, engineering, and commerce, along with technological innovations, has led to an unprecedented growth and enthusiasm in big data collection, thereby redefining computational efficiency and inviting system solutions. For example, the recent AlphaGo system of Deepmind that beats top human Go players needed 1900 CPUs and 280 GPUs to carry out the computation. How to balance computation with communication in this vast distributed cluster, without compromising system throughput or correctness? On the other hand, a small startup developing a mobile app may not afford the same computational power as Google, hence often has to turn into primitive solutions. How to build an algorithmic framework for ML that provides ''knobs'' to adjust the computational load, with explicit, controllable loss on the accuracy? Meeting such diverse computational needs in the big data era has thus been a grand challenge for the ML field.******We attempt to address such computational challenge in ML and big data, through three complementary objectives: (1) Real problems are hard, but also structured. Over the years the importance of designing statistical methodologies and computational algorithms that can exploit certain structure in data and model has become evident. Encouraged by our previous work on sparsity and low-rankness, we propose to investigate two additional structures that are common in ML applications: monotonicity and multi-modality (in the tensor format), and developing efficient algorithms that benefit from the presence of such structures. (2) Data is always noisy and full of random fluctuations, hence diminishing the need of obtaining exact or even high-precision solutions in ML. Approximate computation, if done properly, can significantly reduce the computation time in ML. We initiate a systematic study of the tradeoffs of approximate computation in ML, from ''downgrading'' computationally expensive programs to simpler and cheaper ones, to ''optimally" smooth nondifferentiable functions, and to attach measures of nonconvexity to nonconvex functions. (3) Distributed computation has become the norm in handling big datasets. We propose the Bounded Asynchronous Protocol (BAP) to better balance communication and computation in distributed ML systems, and we continue to investigate the speedups and convergence guarantees of typical ML iterative algorithms under BAP and possibly less stringent convex or smooth assumptions. Our work will further advance the computational theory and practice in ML, and the resulting algorithms and system will be fundamental for analyzing big datasets using ML methodologies.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Foundations of Machine Learning in the Era of Big Data
-
批准号:RGPIN-2017-05032
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$4.08万
-
财政年份:2022
-
负责人:Yu, Yaoliang
-
依托单位:
A Theoretical Foundation and Practical Platform for Adversarial Machine Learning
-
批准号:543522-2019
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.08万
-
财政年份:2021
-
负责人:Yu, Yaoliang
-
依托单位:
Computational Foundations of Machine Learning in the Era of Big Data
-
批准号:RGPIN-2017-05032
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2021
-
负责人:Yu, Yaoliang
-
依托单位:
Computational Foundations of Machine Learning in the Era of Big Data
-
批准号:RGPIN-2017-05032
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2020
-
负责人:Yu, Yaoliang
-
依托单位:
A Theoretical Foundation and Practical Platform for Adversarial Machine Learning
-
批准号:543522-2019
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.08万
-
财政年份:2020
-
负责人:Yu, Yaoliang
-
依托单位:
Computational Foundations of Machine Learning in the Era of Big Data
-
批准号:RGPIN-2017-05032
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2019
-
负责人:Yu, Yaoliang
-
依托单位:
A Theoretical Foundation and Practical Platform for Adversarial Machine Learning
-
批准号:543522-2019
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.08万
-
财政年份:2019
-
负责人:Yu, Yaoliang
-
依托单位:
Computational Foundations of Machine Learning in the Era of Big Data
-
批准号:RGPIN-2017-05032
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2017
-
负责人:Yu, Yaoliang
-
依托单位:
海外基金