Collaborative Research: Improving Techniques of Automatic Speech Recognition and Transfer Learning using Documentary Linguistic Corpora
Collaborative Research: Improving Techniques of Automatic Speech Recognition and Transfer Learning using Documentary Linguistic Corpora
批准号:
2123578
负责人:
Jonathan Amith
金额:
$23.96万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-12-01 至 2025-05-31
中文摘要
越来越多的计算工具,特别是自动语音识别(即语音到文本的转换),被用于促进和调解各行各业的交流。医生对着电脑说话,电脑会把他们的话转录成易读的书面摘要;在线虚拟助手在各种情况下的支持网络中无处不在;终端用户越来越希望他们的语音能够被手机、导航设备和Alexa等工具理解、处理并采取行动。然而,这种机制的建立目前依赖于大量的训练数据(语音和文本),这些数据只能用于主要语言。当只有10小时的转录音频可用时,开发语音识别系统是相当具有挑战性的。解决这一问题的一种方法是通过迁移学习,在这种方法中,语音识别器在一种濒危语言(50小时的转录音频)的相对大量数据上进行训练,然后扩展到相关语言,只需要开发少量的语料库(10小时的转录音频和90小时的未转录音频)。这个项目的目标既是理论的,也是实质性的。首先,本项目将推动低资源语言的自然语言处理的发展,并建立一个将其扩展到其他相关语言的协议。从本质上讲,这个项目将产生一个前所未有的五种相关语言的转录音频语料库,促进理论和描述语言学家对这些语言的比较研究。最先进的自动语音识别(ASR)依赖于材料语料库(带有时间编码转录的音频记录)的存在和人工智能系统的应用,人工智能系统利用神经网络通过解释原始数据来复制人类的学习。这个项目采用了所谓的“端到端神经网络”。有效地,人工神经网络提供输入数据(声学语音信号)和准备的最终结果(转录),并学习实现相同的结果。为了实现这一点,原始语料库被分为训练集(~ 80%)、验证集(~10%)和测试集(~10%)。对于濒危语言的记录,目标不仅仅是ASR系统的准确性,还包括减少人类为实现高度准确的时间编码转录而付出的努力,这些转录将作为目标语言的永久记录存档。项目团队已经为一种语音上困难的声调语言(字符错误率为8%)开发了一个高度准确的系统,并将生成准确的时间编码转录所需的人力减少了75%(从人类从头开始需要40小时到人类校对由ASR生成的转录所需的9小时)。在这个项目中,同一个团队将探索一种形态复杂的粘连语言的ASR策略,希望达到同样程度的准确性,并减少人类的努力。该项目还将解决最先进的ASR面临的另一个挑战:将针对一种语言开发的有效系统转移到资源匮乏、几乎未记录的相关语言。如果该项目取得成功,它将成为其他语言和语言群体类似努力的典范。这些数据和发现将在宾夕法尼亚大学的语言数据联盟和俄克拉何马大学的Sam Noble俄克拉何马自然历史博物馆中提供。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Increasingly, computational tools, particularly automatic speech recognition (that is, the conversion of speech to text), are used to facilitate and mediate communication in all walks of life. Doctors speak into their computers, which transcribe their speech into legible written summaries; online virtual assistants have become ubiquitous in support networks in a wide range of situations; and end users increasingly expect their speech to be understood, processed, and acted upon by cell phones, navigation devices, and tools such as Alexa. The creation of such mechanisms, however, is currently dependent upon a large amount of training data (speech and text) that is only available for major languages. It is quite challenging to develop speech recognition systems when only 10 hours of transcribed audio is available. One way of addressing this problem is through transfer learning, in which a speech recognizer is trained on a relatively large amount of data for one endangered language ( 50 hours of transcribed audio) is then extended to related languages for which only a small corpus of material will be developed (10 hours of transcribed audio and 90 hours of untranscribed audio). The objectives of this project are both theoretical and substantive. For the first, this project will advance the development of natural language processing for low-resource languages and establish a protocol for extending this to other related languages. Substantively, this project will produce an unprecedented corpus of transcribed audio for five related languages, facilitating the comparative study of these languages by theoretical and descriptive linguists. State-of-the-art automatic speech recognition (ASR) depends upon the existence of a corpus of material (audio recordings with time-coded transcriptions) and the application of artificial intelligence systems that utilize neural networks to replicate humans learning by interpreting raw data. This present project employs what is called an "end-to-end neural network." Effectively, the artificial neural network is presented with input data (the acoustic speech signal) and a prepared the end result (a transcription) and learns to achieve the same result. To accomplish this, the original corpus is divided into training (~ 80%), validation (~10%), and test (~10%) sets. For endangered language documentation the goal is not simply accuracy of the ASR system but also the reduction of human effort to achieve highly accurate time-coded transcriptions that will be archived as a permanent record of target language. The project team has already developed a highly accurate system for one phonologically difficult tonal language (character error rate 8%) and reduced the human effort required to produce an accurate time-coded transcription by 75% (from 40 hours needed by a human starting from scratch to 9 hours needed by a human proofing a transcription generated by ASR). For this project the same team will explore ASR strategies for a morphologically complex agglutinative language in the hope of achieving the same degree of accuracy and reduction in human effort. This project will also address another challenge for state-of-the-art ASR: The transfer of an effective system developed for one language to low-resource, virtually undocumented related languages. Should the project be successful it will serve as a model for similar efforts with other languages and language groups. The data and findings will be available at Linguistic Data Consortium at the University of Pennsylvania, and Sam Noble Oklahoma Museum of Natural History, University of Oklahoma.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: RI: Medium: From Acoustic Signal to Morphosyntactic Analysis in One End-to-End Neural System
-
批准号:2211952
-
项目类别:Standard Grant
-
资助金额:$30.19万
-
财政年份:2022
-
负责人:Jonathan Amith
-
依托单位:
A comparative database for biologists, botanists, and linguists
-
批准号:2109821
-
项目类别:Standard Grant
-
资助金额:$41.76万
-
财政年份:2021
-
负责人:Jonathan Amith
-
依托单位:
Documentation of discourse and cultural activities to advance scientific knowledge of an endangered tonal language
-
批准号:1761421
-
项目类别:Continuing Grant
-
资助金额:$17.49万
-
财政年份:2018
-
负责人:Jonathan Amith
-
依托单位:
Collaborative Research: Contributions of Endangered Language Data for Advances in Technology-enhanced Speech Annotation
-
批准号:1500595
-
项目类别:Standard Grant
-
资助金额:$22.78万
-
财政年份:2015
-
负责人:Jonathan Amith
-
依托单位:
Documenting Traditional Ecological Knowledge in the Sierra Nororiental de Puebla, Mexico, in Synchronic and Diachronic Perspectives
-
批准号:1401178
-
项目类别:Standard Grant
-
资助金额:$44.99万
-
财政年份:2014
-
负责人:Jonathan Amith
-
依托单位:
Corpus and lexicon development: Endangered genres of discourse and domains of cultural knowledge in Tu'un isavi (Mixtec) of Yoloxochitl, Guerrero
-
批准号:0966462
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2010
-
负责人:Jonathan Amith
-
依托单位:
Nahuatl Language Documentation Project: Sierra Norte de Puebla [ISO 639 azz]
-
批准号:0756536
-
项目类别:Standard Grant
-
资助金额:$29.18万
-
财政年份:2008
-
负责人:Jonathan Amith
-
依托单位:
Guerrero Nahuatl Language Documentation and Lexicon Enrichment Project
-
批准号:0504164
-
项目类别:Standard Grant
-
资助金额:$29.99万
-
财政年份:2005
-
负责人:Jonathan Amith
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: