Dialect/accent classification via boosted word modeling

Dialect/accent classification via boosted word modeling
复制标题

通过增强词建模进行方言/口音分类

DOI:
10.1109/icassp.2005.1415181
复制
发表时间:
2005
期刊:
Proceedings. (ICASSP '05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005.
影响因子:
--
通讯作者:
J. Hansen
J. Hansen
中科院分区:
--
文献类型:
--
作者:
Rongqing Huang;J. Hansen

文献摘要

被引文献

相似文献

本文介绍了英语方言/口音分类/识别的新进展。提出了一种基于词级的建模技术,其性能优于基于LVCSR的系统,并且具有明显更少的计算代价。该算法将独立于文本的决策问题转化为依赖于文本的问题,并在单词层面产生多个组合决策,而不是在话语层面上进行单一决策。有两组分类器用于WDC,单词分类器,D/SUB W(K)/,以及话语分类器,D/SUB u/。D/SUB W(K)/通过实数AdaBoost.MH算法直接在概率空间而不是特征空间中进行提升。通过单词的方言依存信息来增强D/SUB U/。评估中使用了两个方言语料库。这两个语料库在方言分类方面都取得了显著的进步。
The paper addresses novel advances in English dialect/accent classification/identification. A word level based modeling technique is proposed that is shown to outperform a LVCSR based system with significantly less computational cost. The new algorithm, which is named WDC (word-based dialect classification), converts the text independent decision problem into a text dependent problem and produces multiple combination decisions at the word level rather than make a single decision at the utterance level. There are two sets of classifiers employed for WDC, word classifier, D/sub W(k)/, and utterance classifier, D/sub u/. D/sub W(k)/ is boosted via the real AdaBoost.MH algorithm in the probability space directly instead of the feature space. D/sub u/ is boosted via the dialect dependency information of the words. Two dialect corpora are used in the evaluation. Significant improvement in dialect classification is achieved for both corpora.