课题基金 / 基金详情

Doctoral Dissertation: Investigating the role of grammatical representation in language learnability

Doctoral Dissertation: Investigating the role of grammatical representation in language learnability
博士论文:研究语法表征在语言可学习性中的作用
批准号:
1420785
负责人:
Edward Gibson
金额:
$1.17万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-15 至 2015-12-31

项目摘要

项目成果

Edward Gibson的其他基金

相似基金

相关文献

中文摘要
翻译
处理自然语言的技术在过去十年中变得无处不在。例如,网络搜索引擎处理数十亿页的文本,以确定哪些页面最匹配用户的搜索查询。许多与计算机交互的界面-例如,苹果的Siri个人助理-从用户那里接收语音命令,并且必须处理这些命令以遵循用户的指令。最后,机器翻译技术已经可以用于世界上许多最常见的语言,允许用户自动翻译他们在外国书籍或网站上找到的文本。这些技术主要依赖于简单的语言模型,称为n-gram模型或上下文无关语法,它们在20世纪50年代和60年代开发,并在随后的几十年中得到改进。这些简单的语言模型有很多优点,最值得注意的是,它们可以用来非常快速地处理大量数据。然而,由于它们的简单性,这些模型无法捕捉自然语言中的许多方面的含义。这导致了上述技术的局限性;虚拟个人助理只能处理非常简单的指令类型,机器翻译仍然远不如人工翻译准确。在目前的项目中,莱昂卑尔根和爱德华吉布森博士将研究更复杂的语言模型,目标是提高计算机理解语言的能力。在吉布森博士的指导下,伯根先生将研究被称为轻度上下文敏感语法的语言模型。这些语法能够表达人类拥有的某些类型的语言知识,但这些知识不能用更简单的语法形式来表达。例如,以英语为母语的人知道,像“玛丽踢了球”这样的陈述句在意义上与“玛丽踢了什么?虽然这一事实似乎是显而易见的,但使用简单类型的语法很难(或不可能)表达。然而,轻度上下文敏感的语法可以用一种非常自然的方式来表达这种知识。卑尔根先生和吉布森博士将研究是否可以从语法句子的例子中自动学习轻度上下文敏感的语法。为了做到这一点,他们将使用机器学习技术,这是计算机科学和统计学的一个分支,开发可以自动从数据中学习的算法。研究人员将把这些学习算法与他们的语法形式主义结合起来,并测试他们的方法是否能学习到准确的语法。语法的准确性将使用语料库-句子的集合-进行评估,其中每个句子都用正确的语法结构进行了人工注释。如果可以以这种方式学习准确的适度上下文敏感的语法,那么这为改进上面讨论的自然语言处理技术提供了一种潜在的方法。特别是,因为这种方法不需要专家写下语言的完整语法,所以它有可能在没有巨大工程努力的情况下部署,并且可以很容易地部署在外语中。
英文摘要
Technologies which process natural language have become ubiquitous in the last decade. Web search engines, for example, process billions of pages of text, in order to determine which of those pages best match a user's search query. Many interfaces for interacting with computers -- for example, Apple's Siri personal assistant -- take voice-issued commands from their users, and must process these commands in order to follow the users' instructions. Finally, machine translation technologies have become available for many of the world's most common languages, allowing users to automatically translate text that they find in foreign books or websites. These technologies mostly rely on simple models of language, known as n-gram models or context-free grammars, which were developed in the 1950's and 1960's, and refined in later decades. These simple models of language have many advantages, most notably that they can be used to process large amounts of data very quickly. Because of their simplicity, however, these models are not able to capture many aspects of meaning in natural language. This has resulted in limitations for the technologies discussed above; virtual personal assistants are only able to process very simple types of instructions, and machine translations is still far from being as accurate as human translation. In the current project, Leon Bergen and Dr. Edward Gibson will be investigating more sophisticated kinds of language models, with the goal of increasing the ability of computers to understand language.Under the direction of Dr. Gibson, Mr. Berger will be studying language models known as mildly context-sensitive grammars. These grammars are able to express certain types of linguistic knowledge that humans have, but which cannot be expressed using simpler types of grammatical formalisms. For example, native speakers of English know that a declarative sentence like "Mary kicked the ball" is closely related in meaning to the question "What did Mary kick?" Although this fact seems obvious, it is difficult (or impossible) to express using simple types of grammars. However, mildly context-sensitive grammars can be used to express this knowledge in a very natural way. Mr. Bergen and Dr. Gibson will be studying whether mildly context-sensitive grammars can be automatically learned from examples of grammatical sentences. To do this, they will be using techniques from machine learning, a branch of computer science and statistics that develops algorithms that can automatically learn from data. The researchers will integrate these learning algorithms with their grammatical formalism, and will test whether their method learns an accurate grammar. The accuracy of the grammar will be evaluated using a corpus -- a collection of sentences -- in which every sentence has been manually annotated with its correct grammatical structure. If accurate mildly context-sensitive grammars can be learned in this manner, then this provides a potential method for improving the natural language processing technologies which were discussed above. In particular, because this method does not require an expert to write down the complete grammar for a language, it has the potential to be deployed without tremendous engineering effort, and may be deployed easily in foreign languages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Evaluating meaning-based explanations of syntactic island effects cross-linguistically
Expanding the reach, impact and sustainability of ToyBox Study Malaysia: a kindergarten-based healthy behaviour intervention
  • 批准号:
    MR/V00607X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $6.79万
  • 财政年份:
    2020
  • 负责人:
    Edward Gibson
  • 依托单位:
Improving healthy energy balance- and obesity-related behaviours among preschoolers in Malaysia: feasibility of adapting the ToyBox-Study
  • 批准号:
    MR/P013805/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $39.57万
  • 财政年份:
    2017
  • 负责人:
    Edward Gibson
  • 依托单位:
Workshop on Language Processing and Language Evolution: Special Session at the 2017 CUNY Conference on Human Sentence Processing
海外基金