Learning Representations for Nonspeech Audio Events Through Their Similarities to Speech Patterns

Learning Representations for Nonspeech Audio Events Through Their Similarities to Speech Patterns
复制标题

DOI:
10.1109/taslp.2016.2530401
复制
发表时间:
2016-04-01
影响因子:
5.4
通讯作者:
Mertins, Alfred
Mertins, Alfred
中科院分区:
计算机科学2区
文献类型:
--
作者:
Huy Phan;Hertel, Lars;Mertins, Alfred

文献摘要

被引文献

相似文献

人类听觉系统与人类语音和环境声音非常匹配。因此,出现了这样的问题:人类语音材料是否可以为用于分析非语音音频信号的训练系统提供有用的信息,例如,在分类任务中。为了回答这个问题,我们认为语音模式作为基本的声学概念,体现和代表的目标非语音信号。为了找出非语音信号与语音的相似程度,我们用一个在语音模式上训练的分类器对其进行分类,并使用分类后验来表示与语音库的接近程度。语音相似度最终被用作描述符来表示目标信号。我们进一步表明,一个更好的描述符,可以通过学习组织语音类别层次与树结构。此外,这些描述符是通用的。也就是说,一旦语音分类器已经被学习,它就可以被用作不同数据集的特征提取器而无需再训练。最后,我们提出了一个算法来选择一个足够的子集,它提供了一个近似的表示能力的整个可用的语音模式集。我们对音频事件分析的应用进行了实验。使用TIMIT数据集的音素三元组作为语音模式来学习具有不同复杂度的三个不同数据集(包括UPC-TALP、弗赖堡-106和NAR)的音频事件的描述符。事件分类任务的实验结果表明,即使使用简单的线性分类器,也可以很容易地获得良好的性能。此外,融合学习的描述符作为一个额外的源导致所有三个目标数据集上的最先进的性能。
The human auditory system is very well matched to both human speech and environmental sounds. Therefore, the question arises whether human speech material may provide useful information for training systems for analyzing nonspeech audio signals, e.g., in a classification task. In order to answer this question, we consider speech patterns as basic acoustic concepts, which embody and represent the target nonspeech signal. To find out how similar the nonspeech signal is to speech, we classify it with a classifier trained on the speech patterns and use the classification posteriors to represent the closeness to the speech bases. The speech similarities are finally employed as a descriptor to represent the target signal. We further show that a better descriptor can be obtained by learning to organize the speech categories hierarchically with a tree structure. Furthermore, these descriptors are generic. That is, once the speech classifier has been learned, it can be employed as a feature extractor for different datasets without retraining. Lastly, we propose an algorithm to select a sufficient subset, which provides an approximate representation capability of the entire set of available speech patterns. We conduct experiments for the application of audio event analysis. Phone triplets from the TIMIT dataset were used as speech patterns to learn the descriptors for audio events of three different datasets with different complexity, including UPC-TALP, Freiburg-106, and NAR. The experimental results on the event classification task show that a good performance can be easily obtained even if a simple linear classifier is used. Furthermore, fusion of the learned descriptors as an additional source leads to state-of-the-art performance on all the three target datasets.