Collaborative Research: Deep learning speech recognition for documenting Seneca, a Native American language, and other acutely under-resourced languages
Collaborative Research: Deep learning speech recognition for documenting Seneca, a Native American language, and other acutely under-resourced languages
批准号:
1761477
负责人:
Karin Michelson
金额:
$9.04万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2022-05-31
中文摘要
当探险家雅克·卡地亚在15世纪30年代沿着圣劳伦斯河逆流而上时,六国邦联或Haudenosaunee人的易洛魁语言首次出现。词典和语法是存在的,但每种语言的文档的基本要素也应该包括不同体裁的注释文本。正如美国国会1990年通过的《美洲原住民语言法》所承认的那样,北美土著人民所说的语言具有独特的地位和重要性。该项目将把塞涅卡印第安人民族的成员与一个由语言学和计算机研究人员组成的团队聚集在一起,记录说塞涅卡语的长老们的情况。塞涅卡语是一种特别濒危的易洛魁语。该团队将开发软件,使用自动语音识别(ASR)准确高效地转录这些录音,自动语音识别是Siri或Alexa等数字个人助理背后的技术。Seneca有一个极其复杂的单词结构,称为复合合成,其中一个单词相当于一个从句或句子。这种语言对ASR系统提出了挑战,后者通常被设计为在受限的词汇上识别单词。该项目将通过开发生成合成文本数据的新方法来促进科学知识,以增加对这种复杂性建模所需的现有书面资源。更广泛的影响包括提供用于语言振兴和科学调查的新记录材料。该项目将为来自Seneca Nation的本科生、研究生和年轻人提供宝贵的STEM经验,并扩大美洲原住民在语言和计算机科学方面的参与,包括支持一名Seneca计算机科学博士生。开发的计算工具和方法将供其他致力于记录和分析低资源语言的人使用,这些语言许多是在对国家安全至关重要的地区使用的。Seneca的自发演讲包含长而复杂的单词,但也有许多对理解话语至关重要的短小粒子。塞涅卡语的韵律模式是对塞涅卡语进行切分和注解的关键,这种韵律模式出现在较长的话语中,包括韵律和声调成分。大多数ASR框架将受到多合成形态系统倾向于产生的大词汇量的挑战。此外,ASR系统通常不会对高级韵律信息进行建模。Seneca几乎没有来自自发语音的可用文本数据,这是建立ASR中使用的预测语言模型所必需的,对Seneca学习者来说是无价的。扩大现有的文本数据将需要新的技术来生成合成但可信的文本,特别关注神经序列到序列模型。神经网络模拟远距离和等级关系的能力也将被用来捕捉塞涅卡语中准确分割自发语音所需的发声级韵律模式。该项目汇集了一系列专业知识,并让塞涅卡社区成员、该语言的主要利益攸关方参与进来,将传统的语言方法和计算方法联系起来。通过这种跨学科合作转录和注释的每个新的Seneca录音都将支持Seneca语言的振兴,并有助于推动低资源语言技术的发展。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The Iroquoian languages of the Six Nations Confederacy, or the Haudenosaunee people, were first encountered when the explorer Jacques Cartier sailed up the St. Lawrence River in the 1530s. Dictionaries and grammars exist, but the basic elements of documentation for every language should also include annotated texts of diverse genres. As currently recognized in The Native American Languages Act, passed by the U.S. Congress in 1990, languages spoken by the indigenous peoples of North America have a unique status and importance. This project will bring together members of the Seneca Nation of Indians with a team of linguistics and computing researchers to record elders speaking Seneca, an Iroquoian language that is particularly endangered. The team will develop software to accurately and efficiently transcribe these recordings using automatic speech recognition (ASR), the technology behind digital personal assistants like Siri or Alexa. Seneca has an exceedingly complex word structure, known as polysynthesis, in which a word is equivalent to a clause or sentence. Such languages challenge ASR systems, which are generally designed to recognize words over a constrained vocabulary. This project will advance scientific knowledge by developing novel methods for generating synthetic text data to augment the existing written resources required to model this complexity. Broader impacts include the availability of the newly documented materials for language revitalization and scientific investigation. The project will provide undergraduates, graduate students, and young adults from the Seneca Nation with valuable STEM experience and broadening participation of Native Americans in the language and computing sciences, including supporting a Seneca doctoral student in computer science. The computational tools and methodologies developed will be accessible to others who are working to document and analyze low-resource languages, many spoken in regions of critical importance for national security.Spontaneous speech in Seneca contains long, complex words but also many short particles that are essential to understanding the discourse. Crucial for segmenting and annotating spoken Seneca are the prosodic patterns that occur in longer utterances, involving both metrical and tonal components. Most ASR frameworks would be challenged by the large vocabulary size that a polysynthetic morphological system tends to yield. In addition, ASR systems do not typically model high-level prosodic information. Seneca has little available text data derived from spontaneous speech, which is needed to build the predictive language models used in ASR and is invaluable to Seneca learners. Augmenting the available text data will require novel techniques for generating synthetic but plausible text, with a particular focus on neural sequence-to-sequence models. The ability of neural nets to model long-distance and hierarchical relationships will also be exploited to capture utterance-level prosodic patterns required for accurate segmentation of spontaneous speech in Seneca. By bringing together a range of expertise and by involving Seneca community members, key stakeholders in the language, the project bridges traditional linguistic methodology and computational approaches. Each new Seneca recording that is transcribed and annotated through this collaboration across disciplines will support the revitalization of the Seneca language and help to further the state of the art in low-resource language technology.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
--
发表时间:
2021
期刊:
Language documentation and conservation
影响因子:
1.8
作者:
[Prud'hommeaux, Emily, Jimerson, Robbie, Hatcher, Richard, Michelson, Karin]
通讯作者:
Michelson, Karin
Oneida Prosodic Categories Above the Word
-
批准号:9222382
-
项目类别:Standard Grant
-
资助金额:$4.92万
-
财政年份:1993
-
负责人:Karin Michelson
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: