课题基金 / 基金详情

SGER: Scaling up unsupervised grammar induction

SGER: Scaling up unsupervised grammar induction
SGER:扩大无监督语法归纳
批准号:
0836431
负责人:
Noah Smith
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2009-12-31

项目摘要

项目成果

Noah Smith的其他基金

相似基金

相关文献

中文摘要
翻译
该SGER项目旨在确定计算密集型迭代统计学习算法在MapReduce架构上的可伸缩性。这些算法是自然语言处理中许多研究的基础,但它们对中等规模的训练数据集(文本语料库)的可扩展性一直未得到充分探索。从表面上看,扩展到更多的数据似乎很适合MapReduce范式,这个探索性项目的目的是确定这些算法是否受益于比以前工作中使用的更多数据和更复杂的数据。特别强调了无监督学习算法,如期望最大化算法,这些算法在小问题上得到了广泛的研究,而在大问题上的研究却很少。该技术也适用于许多其他方法。同时,该项目试图探索如何利用超级计算机和MapReduce来使这些学习算法更快,从而允许更快的研究周期。具体地说,“E步”(或其规则)是迭代中计算要求最高的部分,但训练数据独立且相同分布的标准假设允许并行化。在某种程度上,这种并行化受到网络和输入输出开销的影响,每次训练迭代可能会更快,可能会将训练时间从几天或几周减少到几个小时。这个项目探索了这种权衡和其他类似的东西。这项工作利用了雅虎捐赠的资源,供PI的研究小组使用:一台运行Hadoop(MapReduce的开源实现)的4000节点超级计算机。
英文摘要
This SGER project seeks to determine the scalability of computationally intensive, iterative statistical learning algorithms on a MapReduce architecture. Such algorithms underlie much research in natural language processing, yet their scalability to even moderately large training datasets (text corpora) has been under-explored. On the surface, scaling to more data appears to be a good fit for the MapReduce paradigm, and this exploratory project aims to identify whether such algorithms benefit from more data and more complex data than used in prior work. A special emphasis is given to unsupervised learning algorithms, such as the Expectation-Maximization algorithm, which have been widely studied on small problems and rarely studied on large ones. The technique is applicable to many other methods, as well.At the same time, the project seeks to explore how to leverage supercomputers and MapReduce to make these learning algorithms faster, permitting a faster research cycle. Concretely, the "E step" (or itsanalogue) is the most computationally demanding part of an iteration, but the standard assumption that the training data are independently and identically distributed permits parallelization. To the extent that this parallelization is affected by network and input-output overhead, each iteration of training may be made faster, perhaps reducing training time from days or weeks to hours. This project explores this tradeoff and others like it.This work leverages a resource donated by Yahoo for use by the PI's research group: a 4,000-node supercomputer running Hadoop (an open-source implementation of MapReduce).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NSF-BSF: RI: Small: Efficient Transformers via Formal and Empirical Analysis
  • 批准号:
    2113530
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2021
  • 负责人:
    Noah Smith
  • 依托单位:
RI/SES: Conference Proposal: Doctoral Consortium on Text as Data
  • 批准号:
    1830158
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2018
  • 负责人:
    Noah Smith
  • 依托单位:
NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
  • 批准号:
    1813153
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $16.75万
  • 财政年份:
    2018
  • 负责人:
    Noah Smith
  • 依托单位:
RI: Medium: Broad-Coverage Semantic Parsing: Linguistic Representation Learning from Crowd-Scale Data
  • 批准号:
    1562364
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $100.6万
  • 财政年份:
    2016
  • 负责人:
    Noah Smith
  • 依托单位:
海外基金