课题基金 / 基金详情

Parsing and Grammatical Induction as Constraint Solving

Parsing and Grammatical Induction as Constraint Solving
解析和语法归纳作为约束求解
批准号:
RGPIN-2018-06736
负责人:
Dahl, Veronica
金额:
$1.68万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31

项目摘要

项目成果

Dahl, Veronica的其他基金

相似基金

相关文献

中文摘要
翻译
动机:世界上大约有7000种语言,其中仅在加拿大就有60种不同的土著语言。其中许多语言没有得到充分的研究,计算机化和计算机化系统的支持也很差:能够研究这些语言的语言学家数量不足,他们的大部分努力都集中在主流语言上。自动归纳语法的有效方法,可以纳入现代技术和教育工具,将极大地影响语法归纳领域和相关的自然语言处理(NLP)领域,同时显著增加濒危语言的生存机会,从而具有社会经济潜力。******最先进的:大多数语法归纳模型涉及概率(因此学习语法相当于从预先指定的模型族中选择一个模型)或机器学习的统计模型。在精度、可靠性和解释力方面,这些模型不如基于规则和推理的模型,但它们通常可以通过对大量数据集的广泛搜索来克服经典NLP方法的不灵活性,并且当只需要模拟图灵测试意义上的语言理解时,它们可以取得令人印象深刻的结果。一些使用一种语言的语言信息来描述另一种语言的模型,而不是完全解决归纳问题,通常仅限于特定的任务,例如消除另一种语言的歧义,可能需要平行的语料库。******研究在最新技术中的地位:相比之下,在我之前的NSERC资助的支持下,我和我的学生开发了一个从已知语法中归纳未知语法的新模型,它不仅适用于特定任务,也不需要预先指定的模型族,也不需要并行语料库,也不需要任何典型的机器学习模型。我们的模型可以使用该语言的代表性和正确的输入句子以及与该输入相关的词汇来生成未被充分研究的语言的语法,所有这些都是根据已被充分研究的语言的正确语法进行解析的。我们称我们的模型为子宫语法模型(WGM),因为它可以在适当的输入下生成新的语法,就像人类的子宫可以生成所有种族一样。******建议的研究:目前的建议侧重于将我们的实验WGM结果巩固为一个成熟的和可执行的分析和语法推理的计算语言学理论(长期目标),特别关注语义和语用解释(中期)。我们将对这一理论进行微调,将其发展与短期具体的语法归纳(特别是围绕加拿大的语言多样性保护)、解析本身和语言实验(作为测试替代语言约束)的测试平台相结合。
英文摘要
Motivation: Roughly 7,000 languages are spoken in the world, including about 60 distinct indigenous languages in Canada alone. Many of them are under-studied and poorly supported by computational and computerized***systems: the number of linguists who can study them is insufficient and most of their efforts pour into mainstream languages. Efficient methods that automatically induce grammars that can be incorporated into modern technological and educational tools would greatly impact the Grammar Induction field and related Natural Language Processing (NLP) areas, while significantly increasing survival chances for endangered languages, with consequent socio-economic potential.******State-of-the-art: Most grammar induction models involve probabilities (so that learning a grammar amounts to selecting a model from a pre-specified model family) or statistical models of machine learning. With respect to precision, reliability and explanatory power, such models are inferior to rule and inference-based models, but they can often overcome the inflexibilities of classical NLP methods through extensive search on voluminous data sets, and can achieve impressive results when what matters is only to simulate linguistic understanding in the Turing test sense. A few models that use linguistic information from one language for the task of describing another language rather than addressing induction in full, are usually restricted to specific tasks, e.g. disambiguating the other language and may require parallel corpora.******Position of the research within the state-of-the-art: In contrast, with support from my previous NSERC grant, my students and I developed a novel model for inducing unknown grammars from known ones, which works for more than just specific tasks and needs neither a pre-specified model family, nor parallel corpora, nor any of the typical models of machine learning. Our model makes it possible to generate an under-studied language's grammar using representative and correct input sentences in that language, together with its lexicon relevant to that input, all of which is parsed with respect to the correct grammar of a well-studied language. We call our model the Womb Grammar Model (WGM) because it can generate new grammars given appropriate input, much as human wombs can generate all races.******Proposed research: The present proposal focuses on solidifying our experimental WGM results into a full-blown and executable computational linguistic theory of parsing and of grammatical inference (long term goal), with particular focus on semantic and pragmatic interpretation (medium-term). We shall fine-tune this theory by interleaving its development with test-bed short-term concrete applications to grammar induction (in particular, around linguistic diversity preservation in Canada), to parsing per se, and to linguistic experimentation (as testing alternative linguistic constraints).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Parsing and Grammatical Induction as Constraint Solving
  • 批准号:
    RGPIN-2018-06736
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.35万
  • 财政年份:
    2022
  • 负责人:
    Dahl, Veronica
  • 依托单位:
Parsing and Grammatical Induction as Constraint Solving
  • 批准号:
    RGPIN-2018-06736
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2021
  • 负责人:
    Dahl, Veronica
  • 依托单位:
Parsing and Grammatical Induction as Constraint Solving
  • 批准号:
    RGPIN-2018-06736
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2020
  • 负责人:
    Dahl, Veronica
  • 依托单位:
Parsing and Grammatical Induction as Constraint Solving
  • 批准号:
    RGPIN-2018-06736
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.68万
  • 财政年份:
    2019
  • 负责人:
    Dahl, Veronica
  • 依托单位:
海外基金