课题基金 / 基金详情

CAREER: From Rare Events to Competitive Learning Algorithms

CAREER: From Rare Events to Competitive Learning Algorithms
职业:从罕见事件到竞争性学习算法
批准号:
2146334
负责人:
Mesrob Ohannessian
金额:
$54.57万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2027-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
乐观情绪是大量采用数据驱动算法的基础。这样的算法不仅将用于匹配当前的科学和工程解决方案,而且如果有足够的数据,还将用于与未来数年的努力相抗衡。这个项目建议从竞争学习的角度来构建乐观主义,在竞争学习中,一个算法的表现与一系列定制算法一样好,每个算法都有特定的真实性质。因此,即使在不知道真相的情况下,竞争算法也能保证在每种情况下提供最佳可行的性能。这种方法表明,与根据最坏情况来判断业绩的悲观前景相比,还有更多的可能性。它还可以帮助更好地理解(人类)发现的局限性,这是基于自然对被发现的顺应性。通过切入从观察中获取知识的哲学的核心,这项工作将研究与多个层次的数据科学教育努力结合在一起,帮助培养和培训下一代数据科学家和工程师,积极拓展和参与传统上服务不足的学生。该项目充分探索了上下文分布学习的竞争前景,这是一个渗透到数据科学和机器学习中的问题,并将自然语言建模作为特定的用例。这里的目标是演示一种通过从数据中构建随机矩阵或张量来学习预测的算法。如果它的性能可以与一系列了解潜在底层结构(如排名、稀疏性或低流形维度)的丰富专业算法相媲美,那么它就是成功的,尽管没有一个算法是最坏情况下最优的。该项目的更大前提是,增强竞争力的原则是根本性的,涉及多个领域。这些原则包括:(1)退避原则,该原则区分丰富和稀有数据区域,并仅在后者需要的地方适用结构,通过灵活性实现竞争力;(2)经验-贝叶斯原则,它提供了一种机制,在很少有代表性的区域之间共享数据,使它们能够相互帮助;(3)尾部结构,它将前两项原则置于一个易于处理的框架内,并可用来确定基本界限,方法是确定竞争力的必要条件和充分条件。通过严格建立这些原则的基础,该项目的目标是简化竞争性算法的设计,同时发现并适应自然的真相。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
A sense of optimism underlies the mass adoption of data-driven algorithms. Such algorithms will be used not only to match current scientific and engineered solutions but also, given enough data, to rival years of future effort. This project proposes framing optimism in terms of competitive learning, where a single algorithm performs as well as a family of bespoke algorithms, each for a specific true nature. Thus, even without knowledge of the truth, a competitive algorithm is guaranteed to deliver the best feasible performance in each case. This approach shows that much more is possible than the pessimistic outlook of judging performance against a worst-case nature. It can also help better understand the limits of (human) discovery, based on how amenable nature is to being discovered. By cutting right to the heart of the philosophy of acquiring knowledge from observation, this work weaves research with data science education efforts at several levels, to help nurture and train the next generation of data scientists and engineers, with active outreach and engagement of traditionally underserved students.The project fully explores the competitive perspective for contextual distribution learning, a problem that permeates data science and machine learning, with natural language modeling as a specific use case. Here, the goal is to demonstrate an algorithm that learns to predict, by building a stochastic matrix or tensor from data. It is successful if its performance rivals that of a rich array of specialized algorithms aware of potential underlying structures, such as rank, sparsity, or low manifold dimension, despite no single algorithm being worst-case optimal. The larger premise of the project is that the principles that enable competitiveness are fundamental and cut across multiple domains. These include: (i) the back-off principle, which demarcates between abundant and rare data regions and applies structure only in the latter, where it is needed, achieving competitiveness through nimbleness, (ii) the empirical-Bayes principle, which offers a mechanism to share data across rarely represented regions, allowing them to help each other, and (iii) tail structures, which ground the first two principles in a tractable framework and can be used to establish fundamental limits, by characterizing conditions that are necessary and sufficient for competitiveness. By rigorously establishing the foundations of these principles, the goal of the project is to streamline the design of competitive algorithms that simultaneously find the truth of nature and adapt to it.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Rare Metals(稀有金属(英文版))