课题基金 / 基金详情

EAGER: Discovery of Segmental Sub-Word Structure in Speech

EAGER: Discovery of Segmental Sub-Word Structure in Speech
EAGER:语音中分段子词结构的发现
批准号:
1433485
负责人:
Karen Livescu
金额:
$9.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-03-01 至 2015-02-28

项目摘要

项目成果

Karen Livescu的其他基金

相似基金

相关文献

中文摘要
翻译
这一早期的探索性研究概念资助(AGER)研究了新的机器学习技术,以发现语音中的子词单位,用于自动语音识别(ASR)。这一早期探索性研究的概念拨款的表示研究了新的机器学习技术,用于发现用于自动语音识别(ASR)的语音中的子词单元。尽管语言学家在子词单位清单上存在很大的分歧,但很少从数据中学习到用子词单位来表示单词。这个项目代表了一个更大的目标,即在接受有区别的训练的同时,使用科学的见解,使ASR的所有方面都是可学习的,这是探索性工作。与以前的工作不同,语音片段使用有区别地学习的分段相似性来聚类成单元,而不是通过动态时间扭曲或隐马尔可夫模型。学习了多种不同的单位类型,而不是预先假设类似音素的单位。该项目还利用多模式(视频、发音等)数据,通过在多个模式之间共享信息来改进单元发现。在这项探索性工作中,学习的单元被用在一个区分模型中,该模型重新计算了标准基于音素的识别器的初始输出,实验重点是小词汇量/中等词汇量的识别。本项目探索了发现基本语音单元的新方法。除了语音识别的改进之外,该项目还可能对其他研究领域产生广泛影响,这些领域涉及具有分段子结构的序列(如文本、视频、生物数据和金融数据)或涉及聚类。结果还可能包括语言学和言语科学中言语研究的新表现。从社会的角度来看,从长远来看,提高语音识别的可学性将有助于将这项技术更好地移植到服务不足的语言社区,因为这些社区没有大数据集或其他资源的好处。
英文摘要
This EArly Concept Grant for Exploratory Research (EAGER) investigates new machine learning techniques for discovering sub-word units in speech for use in automatic speech recognition (ASR). The representation of this EArly Concept Grant for Exploratory Research investigates new machine learning techniques for discovering sub-word units in speech for use in automatic speech recognition (ASR). The representation of words in terms of sub-word units is rarely learned from data, despite significant disagreement among linguists as to the sub-word unit inventory. This project represents exploratory work toward a larger goal of making all aspects of ASR learnable, using scientific insights while being discriminatively trained.In contrast with prior work, speech segments are clustered into units using discriminatively learned segmental similarities, rather than via dynamic time warping or hidden Markov models. Rather than pre-supposing phoneme-like units, multiple heterogeneous unit typesare learned. The project also leverages multi-modal (video, articulatory, and so on) data to improve unit discovery by sharinginformation across modalities. In this exploratory work, the learned units are used in a discriminative model that rescores initial outputs from a standard phone-based recognizer, and the experiments focus on small-/medium-vocabulary recognition.This project explores new ways of discovering the basic units of speech. Beyond improvements to speech recognition, this project hasthe potential for broad impact on other research areas involving sequences with segmental sub-structure (such as text, video,biological data, and financial data) or involving clustering. The results may also include new representations for the study of speechin linguistics and speech science. From a societal perspective, in the long term making speech recognition more learnable will enableimproved porting of the technology to under-served linguistic communities, which do not have the benefit of large data sets or other resources.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: From acoustics to semantics: Embedding speech for a hierarchy of tasks
RI: Medium: Collaborative Research: Models of Handshape Articulatory Phonology for Recognition and Analysis of American Sign Language
RI: Small: Multi-View Learning of Acoustic Features for Speech Recognition Using Articulatory Measurements
RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
海外基金