Robust Syllable Recognition in the Acousic-Waveform Domain
Robust Syllable Recognition in the Acousic-Waveform Domain
批准号:
EP/D053005/1
负责人:
Zoran Cvetkovic
金额:
$26.44万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2006
资助国家:
英国
项目状态:
已结题
起止时间:
2006 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This proposal is concerned with robust classification/recognition of speech units (phonemes and consonant-vowel syllables) in the domain of acoustic waveforms. The motivation for this research comes from the idea that speech units should be much better separated in the high-dimensional spaces formed by acoustic waveforms than in the smaller representation spaces which are used in state-of-the-art speech recognition systems and which involve significant compression and dimension reduction. Hence, recognition/classification in the acoustic waveform domain should exhibit a higher level of robustness to additive noise than classification in low-dimensional feature spaces.In the first phase of the project we will investigate classification of speech units in the acoustic waveform domain under severe noise conditions, around 0dB signal-to-noise ratio and below, while in the second phase we will study techniques which would make classification robust also to linear filtering. The particular tasks that will be tackled in the first phase can be summarized as follows:1. Study the detailed structure of the sets of acoustic waveforms of individual speech units; in particular their intrinsic dimensions, and the existence of possible nonlinear surfaces on which the data are concentrated.2. Guided by the findings from item 1 above, estimate statistical models of the distribution of speech units in the acoustic waveform domain. We will then design and systematically assess so-called generative classifiers, whose defining property is that they are based on such statistical models.3. Investigate classification of speech units in the acoustic waveform domain using discriminative classification techniques (artificial neural networks, support vector machines, and relevance vector machines). These can be a useful alternative to generative techniques because they focus directly on the classification problem without building explicit models of waveform distributions for each speech unit.4. Construct classifiers by grouping speech units hierarchically. Top-level classifiers will be constructed to distinguish between a small of groups of similar speech units, followed by classifiers separating groups into subgroups and so on. Different methods for defining subgroups will be explored, including confusion matrices of the classifiers from item 3, appropriate distance measures between the statistical models obtained in item 2, and possibly perceptual experiments.A potential argument against our approach is that classification in the acoustic waveform domain will break down in the presence of linear filtering. However, this can be avoided by considering narrow-band signals: for these, the effect of linear filtering is approximately equivalent to amplitude scaling and time delay. In the second phase of the project, we will therefore consider speech classification using narrow-band components of acoustic waveforms. For classification of signals in individual sub-bands, the techniques investigated in the first phase of the project will be considered. A new issue is then how to combine the results of sub-band classifiers to minimize the overall classification error. Here recently developed machine learning techniques will be used, as specified in the case for support.As explained, individual sub-band classifiers should be robust to linear filtering because the latter does not significantly alter the shape of narrow-band signals. On the other hand, the dimension of the spaces of sub-band waveforms will be still high enough to facilitate classification robust to additive noise. Hence, the overall scheme is expected to be robust to both additive noise and linear fitering.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Combined Features and Kernel Design for Noise Robust Phoneme Classification Using Support Vector Machines
使用支持向量机进行噪声稳健音素分类的组合特征和内核设计
DOI:
10.1109/tasl.2010.2090657
发表时间:
2011
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[Yousafzai J]
通讯作者:
Yousafzai J
Towards robust phoneme classification: Augmentation of PLP models with acoustic waveforms
迈向稳健的音素分类:用声学波形增强 PLP 模型
DOI:
--
发表时间:
2008
期刊:
European Signal Processing Conference
影响因子:
--
作者:
[Ager M.]
通讯作者:
Ager M.
Tuning support vector machines for robust phoneme classification with acoustic waveforms
调整支持向量机以利用声学波形进行稳健的音素分类
DOI:
--
发表时间:
2009
期刊:
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
影响因子:
--
作者:
[Yousafzai J.]
通讯作者:
Yousafzai J.
Robust phoneme classification: exploiting the adaptability of acoustic waveform models
鲁棒音素分类:利用声学波形模型的适应性
DOI:
--
发表时间:
期刊:
European Signal Processing Conference, EUSIPCO 2009
影响因子:
--
作者:
[Matthew Ager (Author)]
通讯作者:
Matthew Ager (Author)
Combined PLP - acoustic waveform classification for robust phoneme recognition using support vector machines
组合 PLP - 使用支持向量机进行稳健音素识别的声学波形分类
DOI:
--
发表时间:
2008
期刊:
European Signal Processing Conference, EUSIPCO 2008
影响因子:
--
作者:
[J Yousafzai]
通讯作者:
J Yousafzai
共 7 条
Challenges in Immersive Audio Technology
-
批准号:EP/X032981/1
-
项目类别:Research Grant
-
资助金额:$121.51万
-
财政年份:2024
-
负责人:Zoran Cvetkovic
-
依托单位:
SpeechWave
-
批准号:EP/R012067/1
-
项目类别:Research Grant
-
资助金额:$93.54万
-
财政年份:2018
-
负责人:Zoran Cvetkovic
-
依托单位:
Visits to University of California, Berkeley, Stanford University, and SRI International
-
批准号:EP/K034626/1
-
项目类别:Research Grant
-
资助金额:$2.68万
-
财政年份:2013
-
负责人:Zoran Cvetkovic
-
依托单位:
Perceptual Sound Field Reconstruction and Coherent Emulation
-
批准号:EP/F001142/1
-
项目类别:Research Grant
-
资助金额:$49.67万
-
财政年份:2008
-
负责人:Zoran Cvetkovic
-
依托单位:
海外基金