课题基金 / 基金详情

Collaborative Research: Computational Models for Studying Word Class Distinctions in Polysynthetic Languages

Collaborative Research: Computational Models for Studying Word Class Distinctions in Polysynthetic Languages
协作研究:研究多合成语言中词类区别的计算模型
批准号:
1941742
负责人:
Smaranda Muresan
金额:
$24.93万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-15 至 2023-12-31

项目摘要

项目成果

Smaranda Muresan的其他基金

相似基金

相关文献

中文摘要
翻译
几个世纪以来,大多数语言科学家都同意所有语言都有一些通用的构建模块,名词和动词的类别就是其中之一。所有语言都应该有名词,从“桌子”等具体名词到“冲突”等抽象名词。所有语言都同样应该有动词,比如“去”、“吃”或“爱”。然而,持续的研究认为有些语言没有这些类别。该项目中研究的多合成语言位于该列表的顶部。在多合成语言中,单词由许多部分(语素)组成,这些部分具有独立的含义,但可能无法独立存在。如果多合成语言确实没有名词与动词的区别,那将使它们变得非常不寻常,并且会给我们对认知和言语的普遍原则的理解带来新的挑战。该项目通过开发计算语言学中促进和促进跨语言比较的新方法来探索多合成语言中的名词-动词区别。本研究要解决的具体问题是:(1)是否存在普遍的词类区别,特别是名词和动词之间,如果有,这种区别存在于什么级别? (2)我们能否揭示名词-动词区别的普遍诊断?为了回答这些问题,必须在计算上解决两个关键问题:语态分割和上下文中词汇项的词性标记。研究人员提出了一种基于适配器语法的新型形态分割计算方法,该方法是无监督的,并且能够将语言知识作为归纳偏差包含在内。他们还开发了一种用于词性标记的无监督跨语言迁移方法,该方法将应用于一系列多合成语言。因此,该项目将汇集来自多种多合成语言的计算工具和主要语言数据。除了计算和语言价值外,该项目的结果还将产生重大的社会影响,因为许多多合成语言在对国际安全、语言复兴和健康问题至关重要的领域中使用。对多合成语言的语法和语料库的深入研究将作为新教学材料的基础,用于低资源语言的教学和振兴。该奖项反映了 NSF 的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
For centuries, most language scientists have agreed that all languages have some universal building blocks, and the categories of nouns and verbs are among those. All languages are expected to have nouns, from the concrete ones such as 'table' to the abstract ones such as 'strife'. All languages are equally expected to have verbs, words such as 'go', 'eat', or 'love'. However, a persistent thread of research maintains that there are languages that do not have these categories. Polysynthetic languages investigated in this project are at the top of this list. In a polysynthetic language, words are composed of many parts (morphemes) which have independent meaning but may not be able to stand alone. If polysynthetic languages indeed do not have noun-verb distinctions, that would make them highly unusual and would create new challenges to our understanding of universal principles of cognition and speech. This project explores noun-verb distinctions in polysynthetic languages by developing new methods in computational linguistics that promote and facilitate cross-linguistic comparisons. The specific questions this research addresses are: (1) are there universal word class distinctions, particularly between nouns and verbs, and if yes, at what level does such a distinction exist? (2) can we uncover universal diagnostics for noun-verb distinctions? To answer these questions, two key issues must be addressed computationally: morphological segmentation and part-of-speech tagging of lexical items in context. The researchers propose a novel computational approach to morphological segmentation based on Adaptor Grammars that is unsupervised and is able to include linguistic knowledge as inductive bias. They also develop an unsupervised cross-lingual transfer approach for part-of-speech tagging that will be applied to a range of polysynthetic languages. As a result, the project will assemble computational tools and primary linguistic data from a diverse set of polysynthetic languages. Aside from their computational and linguistic value, the project’s results will also have significant societal impact as many polysynthetic languages are spoken in areas that are key for international security, language revitalization, and health concerns. In-depth work on grammar and corpora of polysynthetic languages will serve as the basis of new pedagogical materials to be used for the teaching and revitalization of low-resource languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Minimally-Supervised Morphological Segmentation using Adaptor Grammars with Linguistic Priors
使用具有语言先验的适配器语法的最小监督形态分割
DOI: 10.18653/v1/2021.findings-acl.347
发表时间: 2021
期刊: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
影响因子: --
作者: [Eskander, Ramy, Lowry, Cass, Khandagale, Sujay, Callejas, Francesca, Klavans, Judith, Polinsky, Maria, Muresan, Smaranda]
通讯作者: Muresan, Smaranda
Unsupervised Stem-based Cross-lingual Part-of-Speech Tagging for Morphologically Rich Low-Resource Languages
针对形态丰富的低资源语言的无监督基于词干的跨语言词性标注
DOI: 10.18653/v1/2022.naacl-main.298
发表时间: 2022
期刊: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
影响因子: --
作者: [Eskander, Ramy, Lowry, Cass, Khandagale, Sujay, Klavans, Judith, Polinsky, Maria, Muresan, Smaranda]
通讯作者: Muresan, Smaranda
EAGER: Collaborative Research:Automated Instruction Assistant for Argumentative Essays
  • 批准号:
    1847853
  • 项目类别:
    Standard Grant
  • 资助金额:
    $14.3万
  • 财政年份:
    2018
  • 负责人:
    Smaranda Muresan
  • 依托单位:
North American Chapter of the Association for Computational Linguistics (NAACL-HLT) 2015 Student Research Workshop
  • 批准号:
    1542303
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2015
  • 负责人:
    Smaranda Muresan
  • 依托单位:
RI: Medium: Collaborative Research: Write A Classifier: Learning Fine-Grained Visual Classifiers from Text and Images
  • 批准号:
    1409257
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $46.32万
  • 财政年份:
    2014
  • 负责人:
    Smaranda Muresan
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)