NSF-BSF: Collaborative Research: RI: Small: Multilingual Language Generation via Understanding of Code Switching
NSF-BSF: Collaborative Research: RI: Small: Multilingual Language Generation via Understanding of Code Switching
批准号:
2007960
负责人:
Yulia Tsvetkov
金额:
$34.56万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2021-11-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Human language technology has recently matured to the extent that computational systems can generally interact with users in ways that are natural to humans, not just to machines. However, most people in the world today are multilingual, and current approaches to language technology do not reflect the reality that multilingual communication is ubiquitous; that is, current technology can interact naturally with monolingual speakers, but not with multilingual ones. Computational systems should be able to generate language that sounds equally natural to these users, and this includes being able to accommodate nonnative speakers. This project first creates a large-scale, broad coverage dataset, reflecting conversations between humans and an automatic system that is sophisticated enough to generate fluent multilingual (i.e. 'code-switched') utterances, but is simple enough for controlled experiments. The dataset is far larger than ones that are currently available, and is based on a much more detailed understanding of language-switching strategies. Second, this dataset is used to develop new methods to incorporate code-switching into contemporary deep-learning language generation, including dialogue systems, question answering, assistive technologies, summarization and machine translation. This innovation should benefit a dramatic number of multilingual computer users, including less privileged users who are currently required to interact with machines in a language they do not speak fluently. Successful completion of the research program will pave the way for the development of natural language technologies that are more accommodating to such users, building bridges over the digital divide. The overarching goal of this project is to develop multilingual and contextualized language generation technologies that are more controllable and more adaptable to multilingual users. The project achieves this goal by completing the following objectives. (1) It develops psycholinguistically-grounded, scalable approaches to collecting corpora for studying how multilingual speakers adapt to each other's linguistic choices in text conversations. These methodologies are employed to collect large-scale, rich datasets of multilingual human-machine conversations. These datasets, as well as additional corpora of human code-switched interactions, should shed new light on the theoretical understanding of cross-lingual usage patterns, allowing for better understanding of how people employ code-switching in written language. (2) It uses the linguistic insights obtained through this endeavor to define classifiers that predict code-switching. (3) Novel approaches are developed for efficient, large-vocabulary neural language generation that incorporate these classifiers, allowing generation systems to introduce code-switching in a way that sounds natural to multilingual users. Consequently, this project should dramatically advance our understanding of code-switching, especially in the relatively unexplored territory of written dialogue. In addition, its contributions benefit a broad range of applications that rely on language generation, including dialogue systems, question answering, assistive technologies, summarization and machine translation.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.18653/v1/2021.findings-acl.338
发表时间:
2021-06
期刊:
Journal of Mathematical Analysis and Applications
影响因子:
1.3
作者:
[Prakhar Gupta;Yulia Tsvetkov;Jeffrey P. Bigham]
通讯作者:
Prakhar Gupta;Yulia Tsvetkov;Jeffrey P. Bigham
DOI:
10.18653/v1/2020.conll-1.46
发表时间:
2020
期刊:
Proceedings of the 24th Conference on Computational Natural Language Learning
影响因子:
--
作者:
[Parekh, Tanmay, Ahn, Emily, Tsvetkov, Yulia, Black, Alan W]
通讯作者:
Black, Alan W
DOI:
10.18653/v1/2021.mrl-1.15
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
作者:
[M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov]
通讯作者:
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov
DOI:
10.18653/v1/2021.naacl-main.240
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
作者:
[Prakhar Gupta;Jeffrey P. Bigham;Yulia Tsvetkov;Amy Pavel]
通讯作者:
Prakhar Gupta;Jeffrey P. Bigham;Yulia Tsvetkov;Amy Pavel
DOI:
--
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
作者:
[Zirui Wang;Yulia Tsvetkov;Orhan Firat;Yuan Cao]
通讯作者:
Zirui Wang;Yulia Tsvetkov;Orhan Firat;Yuan Cao
共 8 条
CAREER: Language Technologies Against the Language of Social Discrimination
-
批准号:2142739
-
项目类别:Continuing Grant
-
资助金额:$55.04万
-
财政年份:2022
-
负责人:Yulia Tsvetkov
-
依托单位:
NSF-BSF: Collaborative Research: RI: Small: Multilingual Language Generation via Understanding of Code Switching
-
批准号:2203097
-
项目类别:Standard Grant
-
资助金额:$34.56万
-
财政年份:2021
-
负责人:Yulia Tsvetkov
-
依托单位:
Collaborative Research: RI: Small: NL(V)P: Natural Language (Variety) Processing
-
批准号:2125201
-
项目类别:Standard Grant
-
资助金额:$16.59万
-
财政年份:2021
-
负责人:Yulia Tsvetkov
-
依托单位:
NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
-
批准号:1812327
-
项目类别:Continuing Grant
-
资助金额:$16.6万
-
财政年份:2018
-
负责人:Yulia Tsvetkov
-
依托单位:
国内基金
海外基金
枯草芽孢杆菌BSF01降解高效氯氰菊酯的种内群体感应机制研究
-
批准号:31871988
-
项目类别:面上项目
-
资助金额:59.0万元
-
批准年份:2018
-
负责人:钟国华
-
依托单位:
基于掺硼直拉单晶硅片的Al-BSF和PERC太阳电池光衰及其抑制的基础研究
-
批准号:61774171
-
项目类别:面上项目
-
资助金额:63.0万元
-
批准年份:2017
-
负责人:艾斌
-
依托单位:
B细胞刺激因子-2(BSF-2)与自身免疫病的关系
-
批准号:38870708
-
项目类别:面上项目
-
资助金额:3.0万元
-
批准年份:1988
-
负责人:吴厚生
-
依托单位: