NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
批准号:
1812778
负责人:
Nathan Schneider
金额:
$16.63万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2021-08-31
中文摘要
当今世界上大多数人都能说多种语言。虽然使用多种语言是一个渐进的现象,但之前的研究主要是针对尚未达到流利程度的第二语言学习者的文本。这个项目的重点是非母语但非常流利的人所写的文本。流利的非母语语言与母语的单语语言在某些概念、结构和搭配的频率上有微妙的不同。这就提出了一种可能性,即语言技术——通常是用“标准”母语训练的——在某些方面存在系统性偏见,使它们对大多数用户的用处不大。该项目将开发方法来检查流利非母语的大型数据集,以检测母语的微妙影响,并为这些语言变体提供自然语言处理(NLP)工具。它的方法将适用于本研究的人群之外,包括基于nlp的社会科学测量和寻求更好地理解双语思维的研究。母语识别将使语言学习、网络安全、地理定位、个性化等方面的潜在应用成为可能。该项目将公开分享实施和数据,并将包括将研究带入教育的教育活动。该项目将推进自然语言处理技术,以揭示不同语言背景的流利使用者在语言使用方面的差异:母语使用者、非常流利的非母语使用者以及翻译人员在将另一种语言翻译成英语时的差异。众所周知,分类器可以通过训练在这些人群中进行高精度的区分,尽管人类很难区分它们。这个项目将专注于语义现象,即使是非母语流利的人也会感到困惑。如果目前的NLP模型偏向于母语,那么它们可能不支持对非母语文本的准确测量;该项目将开发新技术来减轻这种偏见。该项目将提供一系列用于母语识别的新模型,用于语言多样性感知的NLP工具的新测量模型和多品种模型,几种英语的新语义注释,以及对非母语注释的研究。这些研究语言内部变化并将这种变化构建到我们的NLP系统中的新方法将在自然语言语义的计算模型中带来前所未有的灵活性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Most people in the world today are multilingual. Though multilingualism is a gradual phenomenon, previous research has primarily examined text from second language learners who have not yet achieved fluency. This project focuses on text produced by nonnative but highly fluent speakers. Fluent but nonnative language differs subtly from native, monolingual language in the frequencies of certain concepts, constructions, and collocations. This raises the possibility that language technologies -- typically trained on "standard" native language -- are systematically biased in ways that render them less useful for the majority of users. This project will develop methods to examine large datasets of fluent nonnative language to detect the subtle influences of the native language and deliver natural language processing (NLP) tools for these language varieties. Its methods will be applicable beyond the populations in this study, including NLP-based measurement for social science and research seeking to better understand cognition in the bilingual mind. Native language identification will enable potential applications in language learning, cybersecurity, geolocation, personalization, and more. The project will openly share implementations and data, and will include educational activities that bring research into education.This project will advance natural language processing techniques to shed light on the differences in language use by fluent speakers with varying linguistic backgrounds: native speakers, highly fluent nonnative speakers, and translators when translating from another language into English. It is known that classifiers can be trained to discriminate with high accuracy among these populations, even though humans have difficulty telling them apart. This project will focus on semantic phenomena, which can confound even fluent nonnative speakers. If current NLP models are biased toward native language, then they may not support accurate measurement in nonnative text; the project will develop new techniques to mitigate this bias. This project will deliver a range of new models for native language identification, new measurement models and multi-variety models for language-variety-aware NLP tools, new semantic annotations in several Englishes, and a study on nonnative annotation. These novel methods for studying variation within a language and building such variation into our NLP systems will lead to unprecedented flexibility in computational models of natural language semantics.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
作者:
[Siyao Peng;Yang Janet Liu;Yilun Zhu;Austin Blodgett;Yushi Zhao;Nathan Schneider]
通讯作者:
Siyao Peng;Yang Janet Liu;Yilun Zhu;Austin Blodgett;Yushi Zhao;Nathan Schneider
K-SNACS: Annotating Korean Adposition Semantics
K-SNACS:注释韩语介词语义
DOI:
--
发表时间:
2020
期刊:
Proceedings of the Second International Workshop on Designing Meaning Representations
影响因子:
--
作者:
[Hwang, Jena D., Choe, Hanwool, Han, Na-Rae, Schneider, Nathan]
通讯作者:
Schneider, Nathan
DOI:
10.18653/v1/2021.findings-emnlp.423
发表时间:
2021-09
期刊:
影响因子:
--
作者:
[Michael Kranzlein;Nelson F. Liu;Nathan Schneider]
通讯作者:
Michael Kranzlein;Nelson F. Liu;Nathan Schneider
DOI:
10.18653/v1/2021.mwe-1.6
发表时间:
2020-04
期刊:
影响因子:
--
作者:
[Nelson F. Liu;Daniel Hershcovich;Michael Kranzlein;Nathan Schneider]
通讯作者:
Nelson F. Liu;Daniel Hershcovich;Michael Kranzlein;Nathan Schneider
Preparing SNACS for Subjects and Objects
为主体和客体准备 SNACS
DOI:
10.18653/v1/w19-3316
发表时间:
2019
期刊:
Proceedings of the First International Workshop on Designing Meaning Representations
影响因子:
--
作者:
[Shalev, Adi, Hwang, Jena D., Schneider, Nathan, Srikumar, Vivek, Abend, Omri, Rappoport, Ari]
通讯作者:
Rappoport, Ari
共 12 条
CAREER: Metalinguistic Natural Language Understanding
-
批准号:2144881
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2022
-
负责人:Nathan Schneider
-
依托单位:
Collaborative Research: DASS: Transitioning open-source software projects to accountable community governance
-
批准号:2217654
-
项目类别:Standard Grant
-
资助金额:$12.99万
-
财政年份:2022
-
负责人:Nathan Schneider
-
依托单位:
国内基金
海外基金
枯草芽孢杆菌BSF01降解高效氯氰菊酯的种内群体感应机制研究
-
批准号:31871988
-
项目类别:面上项目
-
资助金额:59.0万元
-
批准年份:2018
-
负责人:钟国华
-
依托单位:
基于掺硼直拉单晶硅片的Al-BSF和PERC太阳电池光衰及其抑制的基础研究
-
批准号:61774171
-
项目类别:面上项目
-
资助金额:63.0万元
-
批准年份:2017
-
负责人:艾斌
-
依托单位:
B细胞刺激因子-2(BSF-2)与自身免疫病的关系
-
批准号:38870708
-
项目类别:面上项目
-
资助金额:3.0万元
-
批准年份:1988
-
负责人:吴厚生
-
依托单位: