Neural discovery of abstract inflectional structure
Neural discovery of abstract inflectional structure
批准号:
2217554
负责人:
Micha Elsner
金额:
$33.3万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2026-02-28
中文摘要
在许多语言中,单词根据它们在句子中的语法角色有不同的形式。例如,在标准英语中,当主语是“我/你/我们/他们”和“她/他/它”(“I walk”,“she walks”)时,动词有不同的形式。因为许多语言都有大量不常用的单词和这些单词的许多形式,说话者可能无法记住所有这些形式,必须预测一些。这个项目使用计算技术来衡量单词的语法形式的可预测性,以及单词的哪些信息(例如,它们的声音结构、意义或语法形式的分布)有助于形式的可预测性。该项目通过考虑在可预测的地方出现不可预测形式的情况,改进了以前的技术。例如,英语动词“sing”有不规则过去式“sang”和分词“sung”。如果我们知道动词“give”有一个不规则的过去式“gave”,这(通过类比)也增加了分词不规则的可能性,即使所涉及的不规则形式不同。该项目对世界语言进行了广泛的调查,并对两个语族进行了更集中的研究:罗曼语和闪米特语。该项目有助于科学地理解语言之间的差异,以及所有人类语言在哪些方面必须相似。该项目开发实用的预测系统,可用于提高技术应用程序(例如自动翻译,语音转录应用程序或Alexa等生成流利原始语音的助手)处理罕见单词形式的能力。它还培训了一名与科技行业相关的高级计算技能的研究生,研究人员计划在旨在吸引当地社区的公共活动中讨论该项目。这个项目研究屈折组织如何促进或抑制预测以前未观察到的屈折形式的单词的任务。分布原则在多大程度上塑造了屈折组织,促进了语法形式的预测,以及与形态组织的其他方面的相互作用,在类型学上还没有得到很好的理解。利用形态学中记忆丰富的类比模型的悠久历史,该项目开发了一个计算模型,该模型区分了抽象/分布形态学操作和形态学表面形式,允许独立分析预测每个维度的相对难度。该模型用于进行大规模的类型研究。这个奖项在三个方面影响着社会。首先,本研究的产品有助于为资源不足的语言提供计算工具,对于这些语言,屈折形式预测是具有挑战性的。所有创建的研究软件都是公开的,具有开放访问权限。第二,该项目提供跨学科的STEM培训。第三,本项目支持通过公共活动宣传语言多样性和语言科学研究的重要性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In many languages, words have different forms based on their grammatical role in a sentence. For instance, verbs in Standard English have different forms when the subject is "I/you/we/they" versus "she/he/it" ("I walk", "she walks"). Because many languages have a large number of infrequently used words and many forms of these words, speakers may not be able to memorize all of these forms and must predict some. This project uses computational techniques to measure how predictable the grammatical forms of words are and what information about words (e.g., their sound structure, meaning, or distribution of grammatical forms) contributes to form predictability. The project improves on previous techniques by accounting for cases where unpredictable forms appear in predictable places. For instance, the English verb "sing" has the irregular past tense "sang" and participle "sung". If we know that the verb "give" has an irregular past tense "gave", this (by analogy) increases the chance that the participle is irregular as well, even though the irregular forms involved are different. The project conducts a broad survey of world languages, as well as more focused study of two language families: Romance and Semitic. The project contributes to scientific understanding of the ways in which languages differ from one another and in what respects all human languages must be similar. The project develops practical prediction systems which can be used to improve the ability of technology applications (e.g. automated translators, speech transcription apps, or assistants like Alexa that generate fluent original speech) to handle rare word forms. It also trains a graduate student in advanced computational skills relevant to the technology industry, and the researchers plan to discuss the project in public events designed to engage the local community.This project investigates how inflectional organization facilitates or inhibits the task of predicting previously unobserved inflected forms of words. The extent to which distributional principles shape inflectional organization, facilitate prediction of grammatical forms, and interact with other aspects of morphological organization are typologically not well understood. Drawing on a long history of memory-rich analogical models in morphology, the project develops a computational model which distinguishes between abstract/distributional morphological operations and morphophonological surface form, allowing an independent analysis of the relative difficulty of predicting each dimension. The model is used to conduct a large-scale typological study. This award impacts society in three ways. First, products from this research contribute to computational tools for under-resourced languages, for which inflected form prediction is challenging. All research software created is made publicly available with open-access permissions. Second, the project provides interdisciplinary STEM training. Third, the project supports outreach activities that promote the importance of language diversity and scientific investigation of language via public events.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Analogy in contact: Modeling Maltese plural inflection
接触类比:模拟马耳他语复数变化
DOI:
--
发表时间:
2023
期刊:
Proceedings of the Society for Computation in Linguistics
影响因子:
--
作者:
[Court, Sara, Sims, Andrea D., Elsner, Micha]
通讯作者:
Elsner, Micha
RI: Small: Collaborative Research: Cognitive models of the acquisition of vowels in context
-
批准号:1422987
-
项目类别:Continuing Grant
-
资助金额:$24.0万
-
财政年份:2014
-
负责人:Micha Elsner
-
依托单位:
海外基金