课题基金 / 基金详情

Collaborative Research: Improving Techniques of Automatic Speech Recognition and Transfer Learning using Documentary Linguistic Corpora

Collaborative Research: Improving Techniques of Automatic Speech Recognition and Transfer Learning using Documentary Linguistic Corpora
合作研究:利用文献语言语料库改进自动语音识别和迁移学习技术
批准号:
2123624
负责人:
Shinji Watanabe
金额:
$19.35万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-12-01 至 2025-05-31

项目摘要

项目成果

Shinji Watanabe的其他基金

相似基金

相关文献

中文摘要
翻译
越来越多地使用自动语音识别(即语音到文本的转换)等计算工具来促进和调解通信。医生对着他们的电脑说话,电脑将他们的演讲转录成清晰的书面摘要;在线虚拟助手在各种情况下的支持网络中变得无处不在;最终用户越来越多地希望他们的演讲能够被手机、导航设备和Alexa等工具理解、处理和操作。然而,这种机制的建立目前依赖于大量的训练数据(语音和文本),而这些数据仅适用于主要语文。当只有10个小时的转录音频可用时,开发语音识别系统是相当具有挑战性的。解决这一问题的一种方法是通过迁移学习,其中语音识别器针对一种濒危语言(50小时转录音频)的相对大量的数据进行训练,然后将其扩展到仅为其开发少量语料库的相关语言(10小时转录音频和90小时未转录音频)。该项目的目标既有理论上的,也有实质上的。首先,该项目推动了针对低资源语言的自然语言处理的开发,并建立了将其扩展到其他相关语言的协议。实质上,该项目产生了前所未有的五种相关语言的转录音频语料库,促进了理论和描述性语言学家对这些语言的比较研究。这些数据和发现将在宾夕法尼亚大学的语言数据联盟和俄克拉荷马大学的萨姆·诺布尔俄克拉荷马自然历史博物馆获得。最先进的自动语音识别(ASR)取决于语料库(带有时间编码转录的音频记录)的存在,以及人工智能系统的应用,人工智能系统利用神经网络通过解释原始数据来复制人类的学习。这个项目采用的是所谓的“端到端神经网络”。有效地,向人工神经网络提供输入数据(声学语音信号)和准备的最终结果(转录),并学习以获得相同的结果。为了实现这一点,原始语料库被分为训练集(~80%)、验证集(~10%)和测试集(~10%)。对于濒危语言文件,目标不仅是ASR系统的准确性,而且还减少了人类努力实现高度准确的时间编码抄本,这些抄本将作为目标语言的永久记录存档。该项目团队已经为一种语音困难的声调语言开发了一种高度精确的系统(字符错误率为8%),并将产生准确的时间编码转录所需的人力减少了75%(从人工从头开始需要40小时,到人工校对ASR生成的转录所需的9小时)。对于这个项目,同一个团队将探索一种形态复杂的粘着性语言的ASR策略,希望达到同样的准确度,并减少人类的努力。该项目还将解决最先进的ASR面临的另一个挑战:将为一种语言开发的有效系统转移到资源少、几乎没有文件记录的相关语言。如果该项目成功,它将成为其他语言和语言群体类似努力的典范。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Computational tools such as automatic speech recognition (that is, the conversion of speech to text), are increasingly used to facilitate and mediate communication. Doctors speak into their computers, which transcribe their speech into legible written summaries; online virtual assistants have become ubiquitous in support networks in a wide range of situations; and end users increasingly expect their speech to be understood, processed, and acted upon by cell phones, navigation devices, and tools such as Alexa. The creation of such mechanisms, however, is currently dependent upon a large amount of training data (speech and text) that is only available for major languages. It is quite challenging to develop speech recognition systems when only 10 hours of transcribed audio is available. One way of addressing this problem is through transfer learning, in which a speech recognizer is trained on a relatively large amount of data for one endangered language ( 50 hours of transcribed audio) is then extended to related languages for which only a small corpus of material will be developed (10 hours of transcribed audio and 90 hours of untranscribed audio). The objectives of this project are both theoretical and substantive. For the first, this project advances the development of natural language processing for low-resource languages and establishes a protocol for extending this to other related languages. Substantively, this project produces an unprecedented corpus of transcribed audio for five related languages, facilitating the comparative study of these languages by theoretical and descriptive linguists. The data and findings will be available at Linguistic Data Consortium at the University of Pennsylvania, and Sam Noble Oklahoma Museum of Natural History, University of Oklahoma.State-of-the-art automatic speech recognition (ASR) depends upon the existence of a corpus of material (audio recordings with time-coded transcriptions) and the application of artificial intelligence systems that utilize neural networks to replicate humans learning by interpreting raw data. This present project employs what is called an "end-to-end neural network." Effectively, the artificial neural network is presented with input data (the acoustic speech signal) and a prepared the end result (a transcription) and learns to achieve the same result. To accomplish this, the original corpus is divided into training (~ 80%), validation (~10%), and test (~10%) sets. For endangered language documentation the goal is not simply accuracy of the ASR system but also the reduction of human effort to achieve highly accurate time-coded transcriptions that will be archived as a permanent record of target language. The project team has already developed a highly accurate system for one phonologically difficult tonal language (character error rate 8%) and reduced the human effort required to produce an accurate time-coded transcription by 75% (from 40 hours needed by a human starting from scratch to 9 hours needed by a human proofing a transcription generated by ASR). For this project the same team will explore ASR strategies for a morphologically complex agglutinative language in the hope of achieving the same degree of accuracy and reduction in human effort. This project will also address another challenge for state-of-the-art ASR: The transfer of an effective system developed for one language to low-resource, virtually undocumented related languages. Should the project be successful it will serve as a model for similar efforts with other languages and language groups.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
ML-SUPERB:多语言语音通用性能基准
DOI: 10.21437/interspeech.2023-1316
发表时间: 2023
期刊: ISCA
影响因子: --
作者: [Shi, Jiatong, Berrebbi, Dan, Chen, William, Hu, En-Pei, Huang, Wei-Ping, Chung, Ho-Lam, Chang, Xuankai, Li, Shang-Wen, Mohamed, Abdelrahman, Lee, Hung-yi]
通讯作者: Lee, Hung-yi
Collaborative Research: RI: Medium: Flexible Deep Speech Synthesis through Gestural Modeling
  • 批准号:
    2106929
  • 项目类别:
    Standard Grant
  • 资助金额:
    $39.99万
  • 财政年份:
    2021
  • 负责人:
    Shinji Watanabe
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)