课题基金 / 基金详情

Collaborative Research: Computational Models for Studying Word Class Distinctions in Polysynthetic Languages

Collaborative Research: Computational Models for Studying Word Class Distinctions in Polysynthetic Languages
协作研究:研究多合成语言中词类区别的计算模型
批准号:
1941742
负责人:
Smaranda Muresan
金额:
$24.93万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-15 至 2023-12-31

项目摘要

项目成果

Smaranda Muresan的其他基金

相似基金

相关文献

中文摘要
翻译
几个世纪以来,大多数语言科学家都一致认为,所有语言都有一些通用的构件,名词和动词的类别就是其中之一。所有语言都应该有名词,从具体的名词如“桌子”到抽象的名词如“冲突”。所有的语言都应该有动词,比如“Go”、“Eat”或“Love”。然而,一个持久的研究线索坚持认为,有些语言没有这些类别。本项目中研究的多合成语言位居该列表的首位。在多元综合语言中,单词由许多部分(语素)组成,这些部分具有独立的意义,但可能不能独立存在。如果综合语言确实没有名词和动词的区别,这将使它们非常不寻常,并将给我们对认知和言语的普遍原则的理解带来新的挑战。这个项目通过开发计算语言学中促进和促进跨语言比较的新方法来探索多合成语言中的名词-动词区别。这项研究解决的具体问题是:(1)是否存在普遍的词类差异,特别是名词和动词之间的差异,如果有,这种差异存在于什么水平?(2)我们能否揭示名词-动词差异的普遍诊断?要回答这些问题,必须从计算上解决两个关键问题:词法切分和上下文中词项的词性标注。研究人员提出了一种新的基于Adaptor语法的形态分割的计算方法,该方法是无监督的,能够将语言知识作为归纳偏差包含在内。他们还开发了一种用于词性标注的无监督跨语言迁移方法,该方法将应用于一系列多合成语言。因此,该项目将从一套不同的多种综合语言中收集计算工具和主要语言数据。除了它们的计算和语言价值外,该项目的结果还将产生重大的社会影响,因为许多多语言被用于对国际安全、语言振兴和健康问题至关重要的领域。在综合语言的语法和语料库方面的深入工作将作为新的教学材料的基础,用于教学和振兴低资源语言。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
For centuries, most language scientists have agreed that all languages have some universal building blocks, and the categories of nouns and verbs are among those. All languages are expected to have nouns, from the concrete ones such as 'table' to the abstract ones such as 'strife'. All languages are equally expected to have verbs, words such as 'go', 'eat', or 'love'. However, a persistent thread of research maintains that there are languages that do not have these categories. Polysynthetic languages investigated in this project are at the top of this list. In a polysynthetic language, words are composed of many parts (morphemes) which have independent meaning but may not be able to stand alone. If polysynthetic languages indeed do not have noun-verb distinctions, that would make them highly unusual and would create new challenges to our understanding of universal principles of cognition and speech. This project explores noun-verb distinctions in polysynthetic languages by developing new methods in computational linguistics that promote and facilitate cross-linguistic comparisons. The specific questions this research addresses are: (1) are there universal word class distinctions, particularly between nouns and verbs, and if yes, at what level does such a distinction exist? (2) can we uncover universal diagnostics for noun-verb distinctions? To answer these questions, two key issues must be addressed computationally: morphological segmentation and part-of-speech tagging of lexical items in context. The researchers propose a novel computational approach to morphological segmentation based on Adaptor Grammars that is unsupervised and is able to include linguistic knowledge as inductive bias. They also develop an unsupervised cross-lingual transfer approach for part-of-speech tagging that will be applied to a range of polysynthetic languages. As a result, the project will assemble computational tools and primary linguistic data from a diverse set of polysynthetic languages. Aside from their computational and linguistic value, the project’s results will also have significant societal impact as many polysynthetic languages are spoken in areas that are key for international security, language revitalization, and health concerns. In-depth work on grammar and corpora of polysynthetic languages will serve as the basis of new pedagogical materials to be used for the teaching and revitalization of low-resource languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Minimally-Supervised Morphological Segmentation using Adaptor Grammars with Linguistic Priors
使用具有语言先验的适配器语法的最小监督形态分割
DOI: 10.18653/v1/2021.findings-acl.347
发表时间: 2021
期刊: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
影响因子: --
作者: [Eskander, Ramy, Lowry, Cass, Khandagale, Sujay, Callejas, Francesca, Klavans, Judith, Polinsky, Maria, Muresan, Smaranda]
通讯作者: Muresan, Smaranda
Unsupervised Stem-based Cross-lingual Part-of-Speech Tagging for Morphologically Rich Low-Resource Languages
针对形态丰富的低资源语言的无监督基于词干的跨语言词性标注
DOI: 10.18653/v1/2022.naacl-main.298
发表时间: 2022
期刊: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
影响因子: --
作者: [Eskander, Ramy, Lowry, Cass, Khandagale, Sujay, Klavans, Judith, Polinsky, Maria, Muresan, Smaranda]
通讯作者: Muresan, Smaranda
EAGER: Collaborative Research:Automated Instruction Assistant for Argumentative Essays
  • 批准号:
    1847853
  • 项目类别:
    Standard Grant
  • 资助金额:
    $14.3万
  • 财政年份:
    2018
  • 负责人:
    Smaranda Muresan
  • 依托单位:
North American Chapter of the Association for Computational Linguistics (NAACL-HLT) 2015 Student Research Workshop
  • 批准号:
    1542303
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2015
  • 负责人:
    Smaranda Muresan
  • 依托单位:
RI: Medium: Collaborative Research: Write A Classifier: Learning Fine-Grained Visual Classifiers from Text and Images
  • 批准号:
    1409257
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $46.32万
  • 财政年份:
    2014
  • 负责人:
    Smaranda Muresan
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)