Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM Supervectors
Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM Supervectors
复制标题
融合说话人归一化层次特征和 GMM 超向量的醉酒语音检测
DOI:
10.21437/interspeech.2011-805
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
中科院分区:
文献类型:
--
作者:
Daniel Bone;M. Black;Ming Li;A. Metallinou;Sungbok Lee;Shrikanth S. Narayanan
Speaker state recognition is a challenging problem due to speaker and context variability. Intoxication detection is an important area of paralinguistic speech research with potential real-world applications. In this work, we build upon a base set of various static acoustic features by proposing the combination of several different methods for this learning task. The methods include extracting hierarchical acoustic features, performing iterative speaker normalization, and using a set of GMM supervectors. We obtain an optimal unweighted recall for intoxication recognition using score-level fusion of these subsystems. Unweighted average recall performance is 70.54% on the test set, an improvement of 4.64% absolute (7.04% relative) over the baseline model accuracy of 65.9%. Index Terms: intoxication detection, speaker state, hierarchical features, speaker normalization, GMM supervectors