Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshop

Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshop
复制标题

DOI:
10.1109/icassp.2007.366989
复制
发表时间:
2007-04
期刊:
2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07
影响因子:
--
通讯作者:
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
中科院分区:
其他
文献类型:
--
作者:
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko

文献摘要

被引文献

相似文献

我们报告的调查,在2006年约翰霍普金斯研讨会上进行,到使用发音功能(AF)的观察和语音识别的发音模型。在观察建模领域,我们直接使用AF分类器的输出,在混合HMM/神经网络模型的扩展中,并作为观察向量的一部分,“串联”方法的扩展。在发音建模领域,我们研究了一种模型,该模型具有多个AF状态流,具有软同步约束,用于仅音频和视听识别。该模型实现为动态贝叶斯网络,并测试任务的小词汇量交换机(SVitchboard)语料库和CUAVE视听数字语料库。最后,我们分析AF分类和强制对齐使用一组新收集的功能级手动transmittance。
We report on investigations, conducted at the 2006 Johns Hopkins Workshop, into the use of articulatory features (AFs) for observation and pronunciation models in speech recognition. In the area of observation modeling, we use the outputs of AF classifiers both directly, in an extension of hybrid HMM/neural network models, and as part of the observation vector, an extension of the "tandem" approach. In the area of pronunciation modeling, we investigate a model having multiple streams of AF states with soft synchrony constraints, for both audio-only and audio-visual recognition. The models are implemented as dynamic Bayesian networks, and tested on tasks from the small-vocabulary switchboard (SVitchboard) corpus and the CUAVE audio-visual digits corpus. Finally, we analyze AF classification and forced alignment using a newly collected set of feature-level manual transcriptions.