课题基金 / 基金详情

Discovering and Demonstrating Linguistic Features for Language Documentation

Discovering and Demonstrating Linguistic Features for Language Documentation
发现和展示语言文档的语言特征
批准号:
1761548
负责人:
Graham Neubig
金额:
$45.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-15 至 2023-01-31

项目摘要

项目成果

Graham Neubig的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Documenting endangered languages is matter of great urgency, but is also a time-consuming process. Annotation and curation of speech and text data, searching for interesting, prototypical, or atypical entries, and marshalling examples for pedagogy or publication are still mostly manual processes. This project aims to speed these processes by (1) creating better tools for the automated analysis ofsmaller and partially-annotated corpora, which have potential to reduce the amount of time required by linguists to manually annotate these corpora, and (2) creating better methods for linguists to browse their collected data and answer questions about the characteristics of the language at hand. The intellectual contribution of this proposal will lie in the development of new computational methods for natural language processing (NLP) for endangered languages, and their evaluation, both in controlled environments and as a tool for linguists in the field. It will also have broader impact in the creation of new tools and standards for linguistic documentation,increased collaboration between linguists and computer scientists, and training of a graduate student in the technologies and practices necessary to move this collaboration forward. The training component will increase the STEM workforce capacity in computational linguistics, important given the need for more advanced tools in working on languages that are underdocumented and spoken in countries that are key to national interests. As a specific methodology to realize this vision, this project focuses on recent development of massively multilingual NLP models based on neural networks. These methods work by creating NLP using data from a large number of languages, then using the information gleaned from these languages to improve the accuracy of processing on a new language with a paucity of training data. Within this framework, three major research questions will be examined: (1) How can these techniques be efficiently applied to very-low-resource languages,especially those in the early stages of text collection? (2) What methods can be used to move beyond sentence-by-sentence analyses, and synthesize information about the entirety of the language to propose a simple grammatical specification? (3) Is it possible to provide examples that support typological predictions for a linguist to read and learn more about the nuances of the language they are analyzing? All three of these research questions will be examined in a rigorous process of devising methods, testing on existing data sets for well-resourced languages, and finally deployment to field linguists to examine how they improve the efficiency or accuracy of the language documentation process.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(31)
专著(0)
科研奖励(0)
会议论文
Phoneme Recognition through Fine Tuning of Phonetic Representations: a Case Study on Luhya Language Varieties
通过微调语音表示进行音素识别:Luhya 语言变体的案例研究
DOI: --
发表时间: 2021
期刊: 22nd Annual Conference of the International Speech Communication Association (InterSpeech 2021
影响因子: --
作者: [Siminyu, Kathleen, Li, Xinjian, Anastasopoulos, Antonios, Mortensen, David R., Marlo, Michael, Neubig, Graham]
通讯作者: Neubig, Graham
DOI: --
发表时间: 2020-04
期刊: ArXiv
影响因子: --
作者: [Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig]
通讯作者: Aman Madaan;Shruti Rijhwani;Antonios Anastasopoulos;Yiming Yang;Graham Neubig
DOI: 10.18653/v1/w19-4822
发表时间: 2019-05
期刊:
影响因子: --
作者: [Antonios Anastasopoulos]
通讯作者: Antonios Anastasopoulos
DOI: 10.18653/v1/2020.acl-main.149
发表时间: 2020-05
期刊: ArXiv
影响因子: --
作者: [Emanuele Bugliarello;Sabrina J. Mielke;Antonios Anastasopoulos;Ryan Cotterell;Naoaki Okazaki]
通讯作者: Emanuele Bugliarello;Sabrina J. Mielke;Antonios Anastasopoulos;Ryan Cotterell;Naoaki Okazaki
29
    FAI: Quantifying and Mitigating Disparities in Language Technologies
    • 批准号:
      2040926
    • 项目类别:
      Standard Grant
    • 资助金额:
      $37.5万
    • 财政年份:
      2021
    • 负责人:
      Graham Neubig
    • 依托单位:
    SHF: Small: Open-domain, Data-driven Code Synthesis from Natural Language
    • 批准号:
      1815287
    • 项目类别:
      Standard Grant
    • 资助金额:
      $49.97万
    • 财政年份:
      2018
    • 负责人:
      Graham Neubig
    • 依托单位:
    RI: EAGER: Collaborative Research: Adaptive Heads-up Displays for Simultaneous Interpretation
    • 批准号:
      1748642
    • 项目类别:
      Standard Grant
    • 资助金额:
      $15.0万
    • 财政年份:
      2017
    • 负责人:
      Graham Neubig
    • 依托单位:
    海外基金