CAREER: Modeling Spoken Language Without Parallel Text Annotations
CAREER: Modeling Spoken Language Without Parallel Text Annotations
批准号:
2238605
负责人:
David Harwath
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2028-01-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Automatic speech recognition and understanding technology has been widely adopted into personal digital assistants, automatic transcription of videos and meetings, and many more applications. Building these systems requires massive datasets of speech audio that is human-transcribed into text. When sufficient data is available for a particular domain, modern models based on deep neural networks are capable of highly accurate speech recognition and downstream language understanding tasks. However, for the vast majority of the world's 7,000 languages and even more numerous dialects, large scale annotated datasets simply do not exist, preventing speech technology from serving these languages and their speakers. Inspired by the fact that humans learn to speak long before they can read or write, this CAREER project explores a new paradigm for speech processing that does not rely on transcribed speech. Instead, it develops new models that are capable of learning spoken language directly from speech audio, and applies these models to tasks including building speech recognizers without transcribed speech and automatically translating speech from one language into another. These advances fit within a larger movement in the research community to dramatically reduce the cost and increase the availability of speech recognition and understanding technology to many more languages and users than are served today.This project leverages self-supervised and multimodal learning approaches to automatically discover linguistic structure (phones, words, phrases, etc.) in the raw speech signal which can be treated as ``pseudo-text'' and used in place of conventional text for downstream tasks. It develops new neural network layers for attention-based segmentation of speech, applied in a hierarchical fashion to discover speech units at multiple levels of abstraction. A second novel technique involves adding self-prediction layers and training objectives to a model using the segmentation layers, where the higher layers that would capture word-like structure attempt to predict the tokenization of lower layers that capture sub-word structure. In this way, the model can automatically learn a pronunciation lexicon that captures the compositional relationship between the different tiers of discovered speech units. The project applies these techniques to three downstream applications that are steadily growing in importance in the speech field: unsupervised speech recognition, textless speech-to-speech translation, and textless generation speech for dialog and image captioning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Syllable Discovery and Cross Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model
基于视觉的自我监督语音模型中的音节发现和跨语言泛化
DOI:
--
发表时间:
2023
期刊:
Interspeech
影响因子:
--
作者:
[Peng, Puyuan, Li, Shang-Wen, Rasanen, Okko, Mohamed, Abdelrahman, Harwath, David]
通讯作者:
Harwath, David
DOI:
--
发表时间:
2023
期刊:
IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU
影响因子:
--
作者:
[Lai, Cheng-I Jeff, Shi, Freda, Peng, Puyuan, Kim, Yoon, Gimpel, Kevin, Chang, Shiyu, Chuang, Yung-Sung, Bhati, Saurabhchand, Cox, David, Harwath, David]
通讯作者:
Harwath, David
DOI:
10.48550/arxiv.2305.11095
发表时间:
2023-05
期刊:
影响因子:
--
作者:
[Puyuan Peng;Brian Yan;Shinji Watanabe;David F. Harwath]
通讯作者:
Puyuan Peng;Brian Yan;Shinji Watanabe;David F. Harwath
国内基金
海外基金
Galaxy Analytical Modeling
Evolution (GAME) and cosmological
hydrodynamic simulations.
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2025
-
负责人:Antonios Katsianis
-
依托单位: