Bootstrapping a Corpus of Endangered Languages
Bootstrapping a Corpus of Endangered Languages
批准号:
2319296
负责人:
Emily Prud'hommeaux
金额:
$43.64万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31
中文摘要
理解语言是如何运作的需要对英语以外的语言有更好的理解。迄今为止,大量的研究都集中在英语和一些与之密切相关的语言上。许多悬而未决的问题——甚至是关于英语本身的问题——无法仅用英语和相关语言的数据来回答。这个项目极大地扩展了科学家可以获得的关于16种理论上重要的语言的信息,选择这些语言是因为根据主流理论,它们似乎是不可能的。由于语言是许多人类活动的核心,更好地理解语言可能产生的更广泛影响是巨大的,包括对第二语言教育、语音识别和人工智能等语言技术,以及失语症和阅读障碍等语言相关疾病的康复。该项目的更广泛影响包括支持语言保护和复兴,以及使用16种语言的社区的相关目标。具体来说,这个项目为每种语言生成了大约100万单词的“中等规模”语料库。虽然最近关注的焦点是拥有数十亿单词的大型语料库,但中等规模的语料库在英语和其他高资源语言的计算、心理语言学和习得研究中发挥了关键作用。它们也更可行。该项目采用“自助”方法,首先对所有16种语言的现有材料进行编译、格式化和重新分发,包括纯文本资源和配有转录的音频。然后,它使用尖端的机器学习来开发两种语言的自动语音识别,并评估其在加速新语料库材料转录方面的有用性。然后,这些新材料被用于改进自动语音识别系统,建立一个“良性循环”,加快进一步的工作速度。该方法还可以扩展到其他语言。所有材料和代码都是免费分发的,以刺激研究和工业。该奖项是美国国家科学基金会和美国国家人文基金会为美国国家科学基金会动态语言基础设施- NEH记录濒危语言项目建立的资助伙伴关系的一部分。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding how language works requires a much better understanding of languages other than English. The vast bulk of research to date has focused on English and a few closely-related languages. Many of the outstanding questions – even about English itself - cannot be answered with data only from English and related languages. This project greatly expands the information available to scientists about 16 theoretically-important languages, chosen because they appear to be impossible according to leading theories. Because language is central to much of human activity, potential Broader Impacts of a better understanding of language are vast, including impact on second language education, language technologies such as speech recognition and AI, and rehabilitation of language-related disorders such as aphasia and dyslexia. The Broader Impacts of this project include supporting language preservation and revival as well as related goals of the communities that speak the 16 languages. Specifically, this project produces 'mid-scale' corpora on the order of one million words per language for each language. While much recent focus is on massive corpora with billions of words, mid-scale corpora played a critical role in computational, psycholinguistic, and acquisition studies of English and other high-resource languages. They are also more feasible. This project takes a 'bootstrapping' approach, first compiling, formatting, and redistributing existing materials for all 16 languages, including both text-only resources and audio paired with transcriptions. It then uses cutting edge machine learning to develop Automatic Speech Recognition for two of the languages and assess its usefulness in speeding up transcription of new corpus materials. These new materials are then used to refine the Automatic Speech Recognition, building a 'virtuous cycle' that speeds further work. The method can also be expanded to other languages. All materials and code are distributed for free in order to stimulate research and industry. This award is made as part of a funding partnership between the National Science Foundation and the National Endowment for the Humanities for the NSF Dynamic Language Infrastructure – NEH Documenting Endangered Languages Program.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Deep learning speech recognition for documenting Seneca, a Native American language, and other acutely under-resourced languages
-
批准号:1761562
-
项目类别:Continuing Grant
-
资助金额:$19.6万
-
财政年份:2018
-
负责人:Emily Prud'hommeaux
-
依托单位:
海外基金