CAREER: Leveraging Combinatorial Structures for Robust and Scalable Learning
CAREER: Leveraging Combinatorial Structures for Robust and Scalable Learning
批准号:
1845032
负责人:
Amin Karbasi
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-05-01 至 2025-04-30
中文摘要
为了快速做出明智的决定而搜索大量数据的困难是当今最普遍的挑战之一。许多科学和工程模型以具有固有离散特征的数据为特征,其中离散意味着数据具有有限可能值的集合。这种数据的示例包括文本中的短语到图像中的对象。同样,数据科学的几乎所有方面都涉及离散的任务,如数据总结和模型解释。随着计算方法渗透到科学和工程的各个方面,了解哪些离散公式可以有效地求解以及如何求解是非常重要的。这些问题中的许多都是出了名的难,甚至那些理论上可以解决的问题也可能只适用于少量的数据。然而,实际利益的问题往往表现得更好,并且具有使它们能够更有效地解决的内在结构。该奖项旨在通过开发全新的算法,大幅推进数据科学和机器学习领域的大规模离散优化前沿。该项目还将提供一些教育机会,例如通过耶鲁大学的科学之路项目向当地高中生和中学生提供服务。正如凸性是一个著名的、研究得很好的条件,在这个条件下,连续优化是可处理的,子模块化是一个离散目标可以优化的条件。虽然目前对子模块优化的研究已经导致离散数学规划的根本性突破,但理论与现实世界中从业者使用的现有算法的局限性之间仍然存在很大差距。特别是,大多数现有的子模块优化方法在面对机器学习任务中固有的众多不确定性来源时都失败了,从数据中的噪声到真实目标的可变性。此外,对于各种新颖的机器学习应用来说,子模块化的假设过于强烈,因此需要开发全新的算法。为了将目前可证明的方法从无菌的实验室环境中提升出来,并将其扩展到混乱的现实世界中,重要的是要仔细重新审视它们的局限性,考虑更现实但不太完美的条件,并开发相应的鲁棒但可扩展的算法。这个CAREER项目提出了一个研究计划,旨在设计、分析和评估大规模鲁棒子模块优化的新方法,从而解决一系列具有重要实际意义的优化问题。此外,它还讨论了子模块函数的泛化,这些泛化广泛地扩展了这些方法的适用性,移动到子模块之外的领域。该项目的研究方向具有深远的社会效益,因为在当今信息时代,稳健和可扩展的计算方法在几乎所有科学和工业企业中都发挥着核心作用。这些进展有望在实现数据驱动的科学发现、促进机器学习的公平性以及通过帮助这些社区处理与大数据相关的计算挑战来支持STEM教育方面发挥关键作用。该项目的成果将通过教程、研讨会和开源软件广泛传播给更大的科学界。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The difficulty of searching through a massive amount of data in order to quickly make an informed decision is one of today's most ubiquitous challenges. Many scientific and engineering models feature data with inherently discrete characteristics, where discrete means that the data takes on a finite set of possible values. Examples of such data include phrases in text to objects in an image. Similarly, nearly all aspects of data science involve discrete tasks such as data summarization and model explanation. As computational methods pervade all aspects of science and engineering, it is of great importance to understand which discrete formulations can be solved efficiently and how to do so. Many of these problems are notoriously hard, and even those that are theoretically solvable may only be possible for only small amounts of data. However, the problems of practical interest are often much more well-behaved and possess inherent structure that allows them to be solved more efficiently. This CAREER award aims to substantially advance the frontiers of large-scale discrete optimization in data science and machine learning by developing fundamentally new algorithms. This project will also provide a number of educational opportunities such as outreach to local high school and middle school students through Yale's Pathways to Science program.Just as convexity has been a celebrated and well-studied condition under which continuous optimization is tractable, submodularity is a condition for which discrete objectives may be optimized. While current research in submodular optimization has led to fundamental breakthroughs in discrete mathematical programming, there is still a large gap between the theory and the limitations of the existing algorithms used by practitioners in the real world. In particular, most of the existing submodular optimization methods fail miserably when faced with the numerous sources of uncertainty inherent in machine learning tasks, from noise in the data to variability of the true objective. Moreover, submodularity is too strong an assumption for a variety of novel machine learning applications, necessitating the development of completely new algorithms. In order to lift current provable methods out of the sterile lab environment and scale them into the messy real world, it is important to carefully reexamine their limitations, consider more realistic but less perfect conditions, and develop correspondingly robust yet scalable algorithms. This CAREER project presents a research plan towards designing, analyzing, and evaluating new approaches for robust submodular optimization at a massive scale that leads to solving a broad array of optimization problems of significant practical importance. Furthermore, it addresses generalizations of submodular functions that widely broaden the applicability of these methods, moving to a realm beyond submodularity. The research directions in this project have deep and far-reaching societal benefits, as robust and scalable computational methods play a central role in nearly every scientific and industrial venture in today's information age. Such advances are expected to play crucial roles in enabling data-driven scientific discoveries, promoting fairness in machine learning, and supporting STEM education by helping these communities handle the computational challenges associated with big data. The results of this project will be broadly disseminated to the greater scientific community through tutorials, workshops, and open-source software.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(34)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2021-05
期刊:
影响因子:
--
作者:
[Ji Gao;Amin Karbasi;Mohammad Mahmoody]
通讯作者:
Ji Gao;Amin Karbasi;Mohammad Mahmoody
DOI:
--
发表时间:
2020-02
期刊:
影响因子:
--
作者:
[Yifei Min;Lin Chen;Amin Karbasi]
通讯作者:
Yifei Min;Lin Chen;Amin Karbasi
DOI:
--
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
作者:
[Ruitu Xu;Lin Chen;Amin Karbasi]
通讯作者:
Ruitu Xu;Lin Chen;Amin Karbasi
DOI:
10.48550/arxiv.2207.00486
发表时间:
2022-07
期刊:
影响因子:
--
作者:
[Insu Han;Mike Gartrell;Elvis Dohmatob;Amin Karbasi]
通讯作者:
Insu Han;Mike Gartrell;Elvis Dohmatob;Amin Karbasi
Regret Bounds for Batched Bandits
成批强盗的悔恨界限
DOI:
--
发表时间:
2021
期刊:
Proceedings the AAAI Conference on Human Computation and Crowdsourcing
影响因子:
--
作者:
[Esfandiari, Hossein, Karbasi, Amin, Mehrabian, Abbas, Mirrokni, Vahab]
通讯作者:
Mirrokni, Vahab
共 33 条
海外基金