课题基金 / 基金详情

RI: Small: Multi-View Learning of Acoustic Features for Speech Recognition Using Articulatory Measurements

RI: Small: Multi-View Learning of Acoustic Features for Speech Recognition Using Articulatory Measurements
RI:小:使用发音测量进行语音识别的声学特征的多视图学习
批准号:
1321015
负责人:
Karen Livescu
金额:
$44.49万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-09-01 至 2017-08-31

项目摘要

项目成果

Karen Livescu的其他基金

相似基金

相关文献

中文摘要
翻译
这个项目探讨了学习语音识别的声学特征的技术,基于使用声学和发音记录的多视图学习。 最近的工作表明,通过线性和非线性典型相关分析,使用这种策略的识别改进,其中声学特征的变换被学习,以便最大限度地提高与发音测量(变换)的相关性。 以前的工作一直局限于一个单一的数据库和一种语言。这个项目的主要目标是学习更好的通用功能,为任意扬声器和语言,并开发改进的多视图技术。 项目活动包括:学习时变投影;基于神经网络的多视图技术;使用清晰度、视频、标签等的“多视图”学习;高效的实施;新的输入特征,例如频谱时间滤波器;自动语音识别的一个关键组成部分是音频信号的表示,该音频信号封装了有用的信息,同时丢弃了声学噪声,说话人身份,该项目旨在通过对语音发音器官位置配对的音频记录进行统计分析来自动学习改进的表示(嘴唇、舌头等)和其他测量。 该项目从基本的统计技术开始,并开发新技术,以应对语音和相关信号的挑战和机遇。该项目的影响超出了语音处理。 多视图表示学习的应用包括神经学、气象学、化学计量学、计算机视觉和文本处理;所有这些都可以从改进的技术中受益。 这项工作通过为语音技术课程生成材料以及语音和其他信号的可视化工具来影响教育。
英文摘要
This project explores techniques for learning acoustic features for speech recognition, based on multi-view learning using acoustic and articulatory recordings. Recent work has shown recognition improvements using this strategy via linear and nonlinear canonical correlation analysis, in which transformations of acoustic features are learned so as to maximize correlation with (transformations of) articulatory measurements. Prior work has been limited to a single database and a single language.The main goals of this project are to learn better universal features for arbitrary speakers and languages and to develop improved multi-view techniques. Project activities include: learning time-varying projections; multi-view techniques based on neural networks; "many-view" learning using articulation, video, labels, etc.; efficient implementations; new input features such as spectro-temporal filters; and visualization tools for related research and education.A critical component of automatic speech recognition is a representation of the audio signal that encapsulates useful information while discarding acoustic noise, speaker identity, and so on. This project aims to automatically learn improved representations using statistical analysis of audio recordings paired with positions of the speech articulators (lips, tongue, etc.) and other measurements. The project starts with basic statistical techniques, and develops new techniques that address challenges and opportunities specific to speech and related signals.The project's impact extends beyond speech processing. Applications of multi-view representation learning include neurology, meteorology, chemometrics, computer vision, and text processing; all of these can benefit from the improved techniques. The work impacts education by generating materials for a Speech Technologies course and visualization tools for speech and other signals.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: From acoustics to semantics: Embedding speech for a hierarchy of tasks
EAGER: Discovery of Segmental Sub-Word Structure in Speech
RI: Medium: Collaborative Research: Models of Handshape Articulatory Phonology for Recognition and Analysis of American Sign Language
RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: