Discovering and Demonstrating Linguistic Features for Language Documentation
Discovering and Demonstrating Linguistic Features for Language Documentation
批准号:
1761548
负责人:
Graham Neubig
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-15 至 2023-01-31
中文摘要
记录濒危语言是一件非常紧迫的事情,但也是一个耗时的过程。 语音和文本数据的注释和管理,搜索有趣的,典型的或非典型的条目,以及整理教学或出版的示例仍然主要是手动过程。 该项目旨在通过(1)创建更好的工具来自动分析较小和部分注释的语料库,这有可能减少语言学家手动注释这些语料库所需的时间,以及(2)为语言学家创建更好的方法来浏览他们收集的数据并回答有关手头语言特征的问题。该提案的智力贡献将在于为濒危语言的自然语言处理(NLP)开发新的计算方法,并在受控环境中对其进行评估,并作为该领域语言学家的工具。它还将在创建语言文献的新工具和标准、加强语言学家和计算机科学家之间的合作以及在推动这种合作所需的技术和实践方面培训研究生方面产生更广泛的影响。培训部分将提高STEM工作人员在计算语言学方面的能力,这一点很重要,因为需要更先进的工具来处理对国家利益至关重要的国家中记录不足和使用的语言。作为实现这一愿景的具体方法,该项目重点关注基于神经网络的大规模多语言NLP模型的最新发展。这些方法的工作原理是使用来自大量语言的数据创建NLP,然后使用从这些语言中收集的信息来提高在缺乏训练数据的情况下处理新语言的准确性。在这个框架内,三个主要的研究问题将被检查:(1)这些技术如何有效地应用于非常低的资源的语言,特别是那些在文本收集的早期阶段?(2)什么方法可以用来超越逐句分析,并综合有关整个语言的信息,提出一个简单的语法规范?(3)是否有可能提供支持类型学预测的例子,让语言学家阅读并更多地了解他们正在分析的语言的细微差别?所有这三个研究问题都将在设计方法的严格过程中进行研究,对资源丰富的语言的现有数据集进行测试,最后,授予实地语言学家,以考察他们如何提高语言文档编制过程的效率或准确性。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查进行评估,被认为值得支持的搜索.
英文摘要
Documenting endangered languages is matter of great urgency, but is also a time-consuming process. Annotation and curation of speech and text data, searching for interesting, prototypical, or atypical entries, and marshalling examples for pedagogy or publication are still mostly manual processes. This project aims to speed these processes by (1) creating better tools for the automated analysis ofsmaller and partially-annotated corpora, which have potential to reduce the amount of time required by linguists to manually annotate these corpora, and (2) creating better methods for linguists to browse their collected data and answer questions about the characteristics of the language at hand. The intellectual contribution of this proposal will lie in the development of new computational methods for natural language processing (NLP) for endangered languages, and their evaluation, both in controlled environments and as a tool for linguists in the field. It will also have broader impact in the creation of new tools and standards for linguistic documentation,increased collaboration between linguists and computer scientists, and training of a graduate student in the technologies and practices necessary to move this collaboration forward. The training component will increase the STEM workforce capacity in computational linguistics, important given the need for more advanced tools in working on languages that are underdocumented and spoken in countries that are key to national interests. As a specific methodology to realize this vision, this project focuses on recent development of massively multilingual NLP models based on neural networks. These methods work by creating NLP using data from a large number of languages, then using the information gleaned from these languages to improve the accuracy of processing on a new language with a paucity of training data. Within this framework, three major research questions will be examined: (1) How can these techniques be efficiently applied to very-low-resource languages,especially those in the early stages of text collection? (2) What methods can be used to move beyond sentence-by-sentence analyses, and synthesize information about the entirety of the language to propose a simple grammatical specification? (3) Is it possible to provide examples that support typological predictions for a linguist to read and learn more about the nuances of the language they are analyzing? All three of these research questions will be examined in a rigorous process of devising methods, testing on existing data sets for well-resourced languages, and finally deployment to field linguists to examine how they improve the efficiency or accuracy of the language documentation process.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(31)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Phoneme Recognition through Fine Tuning of Phonetic Representations: a Case Study on Luhya Language Varieties
通过微调语音表示进行音素识别:Luhya 语言变体的案例研究
DOI:
--
发表时间:
2021
期刊:
22nd Annual Conference of the International Speech Communication Association (InterSpeech 2021
影响因子:
--
作者:
[Siminyu, Kathleen, Li, Xinjian, Anastasopoulos, Antonios, Mortensen, David R., Marlo, Michael, Neubig, Graham]
通讯作者:
Neubig, Graham
DOI:
--
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
作者:
[Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig]
通讯作者:
Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
DOI:
10.18653/v1/w19-4822
发表时间:
2019-05
期刊:
影响因子:
--
作者:
[Antonios Anastasopoulos]
通讯作者:
Antonios Anastasopoulos
DOI:
10.18653/v1/2020.acl-main.149
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
作者:
[Emanuele Bugliarello;Sabrina J. Mielke;Antonios Anastasopoulos;Ryan Cotterell;Naoaki Okazaki]
通讯作者:
Emanuele Bugliarello;Sabrina J. Mielke;Antonios Anastasopoulos;Ryan Cotterell;Naoaki Okazaki
DOI:
--
发表时间:
2020
期刊:
and Morphology
影响因子:
--
作者:
[Murikinati, Nikitha, Anastasopoulos, Antonios, Neubig, Graham]
通讯作者:
Neubig, Graham
共 29 条
FAI: Quantifying and Mitigating Disparities in Language Technologies
-
批准号:2040926
-
项目类别:Standard Grant
-
资助金额:$37.5万
-
财政年份:2021
-
负责人:Graham Neubig
-
依托单位:
SHF: Small: Open-domain, Data-driven Code Synthesis from Natural Language
-
批准号:1815287
-
项目类别:Standard Grant
-
资助金额:$49.97万
-
财政年份:2018
-
负责人:Graham Neubig
-
依托单位:
RI: EAGER: Collaborative Research: Adaptive Heads-up Displays for Simultaneous Interpretation
-
批准号:1748642
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2017
-
负责人:Graham Neubig
-
依托单位:
海外基金