课题基金 / 基金详情

EAGER: Discrete Algorithms in NLP

EAGER: Discrete Algorithms in NLP
EAGER:NLP 中的离散算法
批准号:
1451430
负责人:
Hal Daume
金额:
$7.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2015-08-31

项目摘要

项目成果

Hal Daume的其他基金

相似基金

相关文献

中文摘要
翻译
能够理解人类语言的算法必须能够识别该语言的底层结构(例如,主语-动词-宾语)。在自然语言处理社区中开发的计算方法通常已经构建了专门的、一次性的算法,用于解决此类任务中出现的困难的组合优化问题。大多数大型系统都是使用启发式的复杂组合来构建的,这些启发式组合试图使近似搜索技术更好。同时,算法社区已经开发了可扩展的精确算法和近似算法来解决许多这些困难的组合优化问题。这项探索性研究的早期拨款调查了这两个极端之间的联系:语言处理社区需要解决他们需要解决的难题,而算法社区拥有解决这些难题的可证明正确的算法。这一探索解决的最大技术挑战是如何将构建有效语言应用程序所需的统计学习算法与使高效算法成为可能的抽象类型相结合。特别地,这个项目探索了“逆优化”在机器学习中的应用。例如,如果一个人有一个有效的算法来解决一个特定的离散优化问题,他如何学习参数,使这个特定的算法尽可能高的精度?这个项目的成功将产生理论上有原则的、高效的算法来学习解决复杂的语言任务,这些算法可以转化为下游应用,如机器翻译、自动问答和信息检索。该项目的主要技术创新是将“逆优化”问题与在线学习技术相结合。例如,假设最终目标是找到某个特定的结构。对这种结构的搜索通常可以作为一种特殊形式的动态规划问题,而动态规划问题又经常成为超图中的最短路径问题。机器学习的挑战是学习一个模型,在这个模型下,最短路径搜索的解实际上是期望的结构。从算法的角度来看,这需要找到一组输入,在这些输入下给定的结构是最优的:逆优化。然而,给定的结构是最优的是不够的:它还必须以一定的余量击败所有其他(非最优)结构。该项目将开发在线学习算法和逆优化公式的组合,以实现这些进步。
英文摘要
Algorithms that can understand human language must be able to recognize the underlying structure (e.g., subject-verb-object) of that language. Computational approaches developed in the natural language processing community typically have build ad hoc, one-off algorithms for solving the hard, combinatorial optimization problems that arise in such tasks. Most large-scale systems are built using complex combinations of heuristics applied to try to make approximate search techniques better. Concurrently, the algorithms community has developed scalable exact algorithms and approximation algorithms for solving many of these hard combinatorial optimization problems. This EArly Grant for Exploratory Research investigates the connection between these two extremes: the language processing community with the hard problems they need solved, and the algorithms community with the provably correct algorithms for solving such hard problems. The biggest technical challenge this exploration addresses is how to couple the statistical learning algorithms necessary to build effective language applications with the types of abstractions that make efficient algorithms possible. In particular, this project explores the application of "inverse optimization" to machine learning. For example, if one has access to an efficient algorithm for solving a particular discrete optimization problem, how can one learn parameters that make that particular algorithm as high accuracy as possible? Success in this project will give rise to theoretically principled, efficient algorithms for learning to solve complex linguistic tasks, which can transform to downstream applications like machine translation, automatic question answering and information retrieval.This project's main technical innovation is the coupling of "inverse optimization" problems with online learning techniques. For instance, suppose that the end goal is to find some particular structure. The search for this structure can often be cast as a particular form of dynamic programming problem, which in turn often becomes a shortest path problem in a hypergraph. The machine learning challenge then is to learn a model under which the solution to this shortest path search is actually the desired structure. From an algorithmic perspective, this requires finding a set of inputs under which a given structure is optimal: inverse optimization. However, it is not enough for a given structure to be optimal: it must also beat all other (non-optimal) structures by some given margin. This project will develop a combination of online learning algorithms and inverse optimization formulations that enable such advances.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Institute for Trustworthy AI in Law and Society (TRAILS)
  • 批准号:
    2229885
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $2000.0万
  • 财政年份:
    2023
  • 负责人:
    Hal Daume
  • 依托单位:
RI: EAGER: Collaborative Research: Adaptive Heads-up Displays for Simultaneous Interpretation
RI: Small: Linguistic Semantics and Discourse from Leaky Distant Supervision
RI: SMALL: Statistical Linguistic Typology
海外基金