课题基金 / 基金详情

EAGER: Discovery of Segmental Sub-Word Structure in Speech

EAGER: Discovery of Segmental Sub-Word Structure in Speech
EAGER:语音中分段子词结构的发现
批准号:
1433485
负责人:
Karen Livescu
金额:
$9.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-03-01 至 2015-02-28

项目摘要

项目成果

Karen Livescu的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This EArly Concept Grant for Exploratory Research (EAGER) investigates new machine learning techniques for discovering sub-word units in speech for use in automatic speech recognition (ASR). The representation of this EArly Concept Grant for Exploratory Research investigates new machine learning techniques for discovering sub-word units in speech for use in automatic speech recognition (ASR). The representation of words in terms of sub-word units is rarely learned from data, despite significant disagreement among linguists as to the sub-word unit inventory. This project represents exploratory work toward a larger goal of making all aspects of ASR learnable, using scientific insights while being discriminatively trained.In contrast with prior work, speech segments are clustered into units using discriminatively learned segmental similarities, rather than via dynamic time warping or hidden Markov models. Rather than pre-supposing phoneme-like units, multiple heterogeneous unit typesare learned. The project also leverages multi-modal (video, articulatory, and so on) data to improve unit discovery by sharinginformation across modalities. In this exploratory work, the learned units are used in a discriminative model that rescores initial outputs from a standard phone-based recognizer, and the experiments focus on small-/medium-vocabulary recognition.This project explores new ways of discovering the basic units of speech. Beyond improvements to speech recognition, this project hasthe potential for broad impact on other research areas involving sequences with segmental sub-structure (such as text, video,biological data, and financial data) or involving clustering. The results may also include new representations for the study of speechin linguistics and speech science. From a societal perspective, in the long term making speech recognition more learnable will enableimproved porting of the technology to under-served linguistic communities, which do not have the benefit of large data sets or other resources.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: From acoustics to semantics: Embedding speech for a hierarchy of tasks
RI: Medium: Collaborative Research: Models of Handshape Articulatory Phonology for Recognition and Analysis of American Sign Language
RI: Small: Multi-View Learning of Acoustic Features for Speech Recognition Using Articulatory Measurements
RI: Medium: Collaborative Research: Explicit Articulatory Models of Spoken Language, with Application to Automatic Speech Recognition
海外基金