课题基金 / 基金详情

Doctoral Dissertation: Investigating the role of grammatical representation in language learnability

Doctoral Dissertation: Investigating the role of grammatical representation in language learnability
博士论文:研究语法表征在语言可学习性中的作用
批准号:
1420785
负责人:
Edward Gibson
金额:
$1.17万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-15 至 2015-12-31

项目摘要

项目成果

Edward Gibson的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年里,处理自然语言的技术变得无处不在。例如,网络搜索引擎处理数十亿页的文本,以便确定这些页面中的哪些与用户的搜索查询最匹配。许多用于与计算机交互的界面--例如苹果的Siri个人助理--都接受用户通过语音发出的命令,并且必须处理这些命令才能遵循用户的指令。最后,机器翻译技术已经适用于世界上许多最常见的语言,允许用户自动翻译他们在外国书籍或网站上找到的文本。这些技术主要依赖于简单的语言模型,即所谓的n元语法模型或上下文无关文法,它们是在1950年的S和1960年的S中发展起来的,并在后来的几十年里得到了改进。这些简单的语言模型有许多优点,最值得注意的是,它们可以非常快地用于处理大量数据。然而,由于它们的简单性,这些模型无法捕捉到自然语言中意义的许多方面。这导致了上述技术的局限性;虚拟个人助理只能处理非常简单的指令类型,机器翻译仍然远远不像人工翻译那样准确。在目前的项目中,里昂·卑尔根和爱德华·吉布森博士将研究更复杂的语言模型,目标是提高计算机理解语言的能力。在吉布森博士的指导下,伯杰将研究被称为温和语境敏感语法的语言模型。这些语法能够表达人类拥有的某些类型的语言知识,但不能使用更简单类型的语法形式化来表达。例如,以英语为母语的人知道,一个陈述句,如“玛丽踢球”,在意义上与疑问句“玛丽踢了什么?”密切相关。尽管这一事实似乎显而易见,但它很难(或不可能)用简单类型的语法来表达。但是,可以使用温和的上下文敏感语法以一种非常自然的方式来表达这些知识。卑尔根和吉布森博士将研究是否可以从语法句子的例子中自动学习适度的上下文敏感语法。为了做到这一点,他们将使用机器学习的技术,机器学习是计算机科学和统计学的一个分支,开发可以从数据中自动学习的算法。研究人员将把这些学习算法与他们的语法形式主义结合起来,并测试他们的方法是否学习了准确的语法。将使用语料库--一组句子--对语法的准确性进行评估,在语料库中,每个句子都已用正确的语法结构进行了手动注释。如果能够以这种方式学习准确的、温和的上下文敏感语法,则这为改进上面讨论的自然语言处理技术提供了一种潜在的方法。特别是,因为这种方法不需要专家写下一种语言的完整语法,所以它有可能在不需要巨大的工程工作的情况下进行部署,并且可以很容易地在外语中部署。
英文摘要
Technologies which process natural language have become ubiquitous in the last decade. Web search engines, for example, process billions of pages of text, in order to determine which of those pages best match a user's search query. Many interfaces for interacting with computers -- for example, Apple's Siri personal assistant -- take voice-issued commands from their users, and must process these commands in order to follow the users' instructions. Finally, machine translation technologies have become available for many of the world's most common languages, allowing users to automatically translate text that they find in foreign books or websites. These technologies mostly rely on simple models of language, known as n-gram models or context-free grammars, which were developed in the 1950's and 1960's, and refined in later decades. These simple models of language have many advantages, most notably that they can be used to process large amounts of data very quickly. Because of their simplicity, however, these models are not able to capture many aspects of meaning in natural language. This has resulted in limitations for the technologies discussed above; virtual personal assistants are only able to process very simple types of instructions, and machine translations is still far from being as accurate as human translation. In the current project, Leon Bergen and Dr. Edward Gibson will be investigating more sophisticated kinds of language models, with the goal of increasing the ability of computers to understand language.Under the direction of Dr. Gibson, Mr. Berger will be studying language models known as mildly context-sensitive grammars. These grammars are able to express certain types of linguistic knowledge that humans have, but which cannot be expressed using simpler types of grammatical formalisms. For example, native speakers of English know that a declarative sentence like "Mary kicked the ball" is closely related in meaning to the question "What did Mary kick?" Although this fact seems obvious, it is difficult (or impossible) to express using simple types of grammars. However, mildly context-sensitive grammars can be used to express this knowledge in a very natural way. Mr. Bergen and Dr. Gibson will be studying whether mildly context-sensitive grammars can be automatically learned from examples of grammatical sentences. To do this, they will be using techniques from machine learning, a branch of computer science and statistics that develops algorithms that can automatically learn from data. The researchers will integrate these learning algorithms with their grammatical formalism, and will test whether their method learns an accurate grammar. The accuracy of the grammar will be evaluated using a corpus -- a collection of sentences -- in which every sentence has been manually annotated with its correct grammatical structure. If accurate mildly context-sensitive grammars can be learned in this manner, then this provides a potential method for improving the natural language processing technologies which were discussed above. In particular, because this method does not require an expert to write down the complete grammar for a language, it has the potential to be deployed without tremendous engineering effort, and may be deployed easily in foreign languages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Evaluating meaning-based explanations of syntactic island effects cross-linguistically
Expanding the reach, impact and sustainability of ToyBox Study Malaysia: a kindergarten-based healthy behaviour intervention
  • 批准号:
    MR/V00607X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $6.79万
  • 财政年份:
    2020
  • 负责人:
    Edward Gibson
  • 依托单位:
Improving healthy energy balance- and obesity-related behaviours among preschoolers in Malaysia: feasibility of adapting the ToyBox-Study
  • 批准号:
    MR/P013805/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $39.57万
  • 财政年份:
    2017
  • 负责人:
    Edward Gibson
  • 依托单位:
Workshop on Language Processing and Language Evolution: Special Session at the 2017 CUNY Conference on Human Sentence Processing
海外基金