Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM Supervectors

Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM Supervectors
复制标题

融合说话人归一化层次特征和 GMM 超向量的醉酒语音检测

DOI:
10.21437/interspeech.2011-805
复制
发表时间:
2011
期刊:
IEEE International Conference on Multimedia and Expo, 2001. ICME 2001.
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
--
文献类型:
--
作者:
Daniel Bone;M. Black;Ming Li;A. Metallinou;Sungbok Lee;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

由于说话人和语境的多变性,说话人状态识别是一个具有挑战性的问题。醉酒检测是副语言语音研究的一个重要领域,具有潜在的实际应用价值。在这项工作中,我们通过提出几种不同方法的组合来实现这一学习任务,从而建立在各种静态声学特征的基础上。这些方法包括提取分层声学特征,执行迭代说话人归一化,以及使用一组GMM超向量。我们使用这些子系统的分数级融合来获得用于醉酒识别的最优未加权召回率。非加权平均召回性能在测试集上为70.54%,比基准模型65.9%的绝对准确率(7.04%)提高了4.64%。索引词:醉酒检测、说话人状态、层次特征、说话人归一化、GMM超向量
Speaker state recognition is a challenging problem due to speaker and context variability. Intoxication detection is an important area of paralinguistic speech research with potential real-world applications. In this work, we build upon a base set of various static acoustic features by proposing the combination of several different methods for this learning task. The methods include extracting hierarchical acoustic features, performing iterative speaker normalization, and using a set of GMM supervectors. We obtain an optimal unweighted recall for intoxication recognition using score-level fusion of these subsystems. Unweighted average recall performance is 70.54% on the test set, an improvement of 4.64% absolute (7.04% relative) over the baseline model accuracy of 65.9%. Index Terms: intoxication detection, speaker state, hierarchical features, speaker normalization, GMM supervectors