课题基金 / 基金详情

Bootstrapping a Corpus of Endangered Languages

Bootstrapping a Corpus of Endangered Languages
引导濒危语言语料库
批准号:
2319296
负责人:
Emily Prud'hommeaux
金额:
$43.64万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Emily Prud'hommeaux的其他基金

相似基金

相关文献

中文摘要
翻译
要了解语言的工作原理,需要更好地理解英语以外的其他语言。到目前为止,大部分研究都集中在英语和一些密切相关的语言上。许多悬而未决的问题--甚至是关于英语本身--不能仅用英语和相关语言的数据来回答。这个项目极大地扩大了科学家可以获得的关于16种理论上重要的语言的信息,之所以选择这些语言,是因为根据领先的理论,这些语言似乎是不可能的。由于语言是人类活动的核心,更好地理解语言的潜在更广泛的影响是巨大的,包括对第二语言教育、语音识别和人工智能等语言技术以及失语症和阅读障碍等语言相关疾病的康复的影响。该项目的更广泛影响包括支持语言的保存和复兴以及说这16种语言的社区的相关目标。具体地说,该项目为每种语言制作了大约100万字的“中等规模”语料库。虽然最近的重点是拥有数十亿单词的大型语料库,但中等规模的语料库在英语和其他高资源语言的计算、心理语言学和习得研究中发挥了关键作用。它们也更具可行性。该项目采用“自举”方法,首先汇编、格式化和重新分发所有16种语言的现有材料,包括纯文本资源和配对转录的音频。然后,它使用尖端机器学习来开发两种语言的自动语音识别,并评估其在加快新语料库材料转录方面的有用性。然后,这些新材料被用来改进自动语音识别,建立一个加速进一步工作的“良性循环”。该方法还可以扩展到其他语言。所有材料和代码都是免费分发的,以刺激研究和工业。这一奖项是国家科学基金会和国家人文基金会为国家科学基金会动态语言基础设施-NEH记录濒危语言项目建立的资金合作伙伴关系的一部分。该奖项反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding how language works requires a much better understanding of languages other than English. The vast bulk of research to date has focused on English and a few closely-related languages. Many of the outstanding questions – even about English itself - cannot be answered with data only from English and related languages. This project greatly expands the information available to scientists about 16 theoretically-important languages, chosen because they appear to be impossible according to leading theories. Because language is central to much of human activity, potential Broader Impacts of a better understanding of language are vast, including impact on second language education, language technologies such as speech recognition and AI, and rehabilitation of language-related disorders such as aphasia and dyslexia. The Broader Impacts of this project include supporting language preservation and revival as well as related goals of the communities that speak the 16 languages. Specifically, this project produces 'mid-scale' corpora on the order of one million words per language for each language. While much recent focus is on massive corpora with billions of words, mid-scale corpora played a critical role in computational, psycholinguistic, and acquisition studies of English and other high-resource languages. They are also more feasible. This project takes a 'bootstrapping' approach, first compiling, formatting, and redistributing existing materials for all 16 languages, including both text-only resources and audio paired with transcriptions. It then uses cutting edge machine learning to develop Automatic Speech Recognition for two of the languages and assess its usefulness in speeding up transcription of new corpus materials. These new materials are then used to refine the Automatic Speech Recognition, building a 'virtuous cycle' that speeds further work. The method can also be expanded to other languages. All materials and code are distributed for free in order to stimulate research and industry. This award is made as part of a funding partnership between the National Science Foundation and the National Endowment for the Humanities for the NSF Dynamic Language Infrastructure – NEH Documenting Endangered Languages Program.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Deep learning speech recognition for documenting Seneca, a Native American language, and other acutely under-resourced languages
  • 批准号:
    1761562
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $19.6万
  • 财政年份:
    2018
  • 负责人:
    Emily Prud'hommeaux
  • 依托单位:
海外基金