课题基金 / 基金详情

CAREER: Modeling Spoken Language Without Parallel Text Annotations

CAREER: Modeling Spoken Language Without Parallel Text Annotations
职业:在没有并行文本注释的情况下对口语进行建模
批准号:
2238605
负责人:
David Harwath
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2028-01-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
语音自动识别和理解技术已被广泛应用于个人数字助理、视频和会议的自动转录等许多应用中。建立这些系统需要大量的语音音频数据集,这些数据集是人工转录成文本的。当一个特定领域有足够的数据可用时,基于深度神经网络的现代模型能够高度精确的语音识别和下游语言理解任务。然而,对于世界上绝大多数的7000种语言和更多的方言来说,大规模的注释数据集根本不存在,这阻碍了语音技术为这些语言及其使用者提供服务。受人类早在会读或写之前就学会说话这一事实的启发,这个CAREER项目探索了一种不依赖于转录语音的语音处理新范式。相反,它开发了能够直接从语音音频中学习口语的新模型,并将这些模型应用于包括构建语音识别器在内的任务,这些任务不需要转录语音,也可以自动将语音从一种语言翻译成另一种语言。这些进步符合研究界的一个更大的运动,即大幅降低成本,增加语音识别和理解技术的可用性,使其能够为更多的语言和用户提供服务。本项目利用自监督和多模态学习方法,在原始语音信号中自动发现语言结构(电话、单词、短语等),这些语言结构可以被视为“伪文本”,用于替代传统文本进行下游任务。它为基于注意力的语音分割开发了新的神经网络层,以分层的方式应用于在多个抽象层次上发现语音单元。第二种新技术涉及使用分割层向模型添加自预测层和训练目标,其中捕获类词结构的较高层试图预测捕获子词结构的较低层的标记化。通过这种方式,该模型可以自动学习一个发音词典,该词典可以捕获所发现语音单元的不同层之间的组成关系。该项目将这些技术应用于语音领域中重要性稳步增长的三个下游应用:无监督语音识别、无文本语音到语音翻译以及用于对话和图像字幕的无文本生成语音。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automatic speech recognition and understanding technology has been widely adopted into personal digital assistants, automatic transcription of videos and meetings, and many more applications. Building these systems requires massive datasets of speech audio that is human-transcribed into text. When sufficient data is available for a particular domain, modern models based on deep neural networks are capable of highly accurate speech recognition and downstream language understanding tasks. However, for the vast majority of the world's 7,000 languages and even more numerous dialects, large scale annotated datasets simply do not exist, preventing speech technology from serving these languages and their speakers. Inspired by the fact that humans learn to speak long before they can read or write, this CAREER project explores a new paradigm for speech processing that does not rely on transcribed speech. Instead, it develops new models that are capable of learning spoken language directly from speech audio, and applies these models to tasks including building speech recognizers without transcribed speech and automatically translating speech from one language into another. These advances fit within a larger movement in the research community to dramatically reduce the cost and increase the availability of speech recognition and understanding technology to many more languages and users than are served today.This project leverages self-supervised and multimodal learning approaches to automatically discover linguistic structure (phones, words, phrases, etc.) in the raw speech signal which can be treated as ``pseudo-text'' and used in place of conventional text for downstream tasks. It develops new neural network layers for attention-based segmentation of speech, applied in a hierarchical fashion to discover speech units at multiple levels of abstraction. A second novel technique involves adding self-prediction layers and training objectives to a model using the segmentation layers, where the higher layers that would capture word-like structure attempt to predict the tokenization of lower layers that capture sub-word structure. In this way, the model can automatically learn a pronunciation lexicon that captures the compositional relationship between the different tiers of discovered speech units. The project applies these techniques to three downstream applications that are steadily growing in importance in the speech field: unsupervised speech recognition, textless speech-to-speech translation, and textless generation speech for dialog and image captioning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Syllable Discovery and Cross Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model
基于视觉的自我监督语音模型中的音节发现和跨语言泛化
DOI: --
发表时间: 2023
期刊: Interspeech
影响因子: --
作者: [Peng, Puyuan, Li, Shang-Wen, Rasanen, Okko, Mohamed, Abdelrahman, Harwath, David]
通讯作者: Harwath, David
Audio-Visual Neural Syntax Acquisition
视听神经语法获取
DOI: --
发表时间: 2023
期刊: IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU
影响因子: --
作者: [Lai, Cheng-I Jeff, Shi, Freda, Peng, Puyuan, Kim, Yoon, Gimpel, Kevin, Chang, Shiyu, Chuang, Yung-Sung, Bhati, Saurabhchand, Cox, David, Harwath, David]
通讯作者: Harwath, David
DOI: 10.48550/arxiv.2305.11095
发表时间: 2023-05
期刊:
影响因子: --
作者: [Puyuan Peng;Brian Yan;Shinji Watanabe;David F. Harwath]
通讯作者: Puyuan Peng;Brian Yan;Shinji Watanabe;David F. Harwath
国内基金
海外基金
Galaxy Analytical Modeling Evolution (GAME) and cosmological hydrodynamic simulations.
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2025
  • 负责人:
    Antonios Katsianis
  • 依托单位: