Content analysis for audio classification and segmentation

Content analysis for audio classification and segmentation
复制标题

DOI:
10.1109/tsa.2002.804546
复制
发表时间:
2002-10-01
期刊:
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子:
--
通讯作者:
Jiang, H
Jiang, H
中科院分区:
其他
文献类型:
--
作者:
Lu, L;Zhang, HJ;Jiang, H

文献摘要

被引文献

相似文献

In this paper, we present our study of audio content analysis for classification and segmentation, in which an audio stream is segmented according to audio type or speaker identity. We propose a robust approach that is capable of classifying and segmenting an audio stream into speech, music, environment sound, and silence. Audio classification is processed in two steps, which makes it suitable for different applications. The first step of the classification is speech and nonspeech discrimination. In this step, a novel algorithm based on K-nearest-neighbor (KNN) and linear spectral pairs-vector quantization (LSP-VQ) is developed. The second step further divides nonspeech class into music, environment sounds, and silence with a rule-based classification scheme. A set of new features such as the noise frame ratio and band periodicity are introduced and discussed in detail. We also develop an unsupervised speaker segmentation algorithm using a novel scheme based on quasi-GMM and LSP correlation analysis. Without a priori knowledge, this algorithm can support the open-set speaker, online speaker modeling and real time segmentation. Experimental results indicate that the proposed algorithms can produce very satisfactory results.