Online acoustic scene analysis based on nonparametric Bayesian model

Online acoustic scene analysis based on nonparametric Bayesian model
复制标题

DOI:
10.1109/eusipco.2016.7760396
复制
发表时间:
2016-08
期刊:
2016 24th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
Keisuke Imoto;Nobutaka Ono
Keisuke Imoto;Nobutaka Ono
中科院分区:
其他
文献类型:
--
作者:
Keisuke Imoto;Nobutaka Ono

文献摘要

相似文献

在本文中,我们提出了一种新的在线方法,从顺序获得的声音分析声学场景。一种用于分析声学场景的前瞻性方法是使用声学主题和观察到的声音中的事件序列的生成模型,其中声学主题表示将声学场景和声学事件相关联的声学事件的潜在结构。这种生成模型被称为声学主题模型(ATM)。然而,传统的ATM采用批处理技术来估计模型参数,并且不能对顺序获得的声事件序列进行建模。此外,在观测声学事件之前,需要预先确定声学事件序列中声学主题的类别数量。然而,用于表示声学场景的声学主题的必要数量根据其内容而变化,并且这导致声学主题的类别的实际数量与类别的预定数量之间的失配。在我们的方法中,类的声学主题的数量可以自动推断从顺序获得的声学事件序列的基础上的在线和非参数贝叶斯技术。利用真实声音进行的在线声场景估计实验结果表明,该方法的声场景分类性能优于传统的ATM。此外,所提出的方法产生了一个有效的计算性能。
In this paper, we propose a novel online method for analyzing acoustic scenes from sequentially obtained sounds. One prospective method for analyzing acoustic scenes is the use of a generative model of acoustic topics and event sequences in observed sounds, where the acoustic topic represents the latent structure of acoustic events associating an acoustic scene and acoustic events. This generative model is called an acoustic topic model (ATM). However, the conventional ATM employs a batch technique for estimating model parameters and cannot model sequentially obtained acoustic event sequences. Moreover, the number of classes of acoustic topics that lies in acoustic event sequences needs to be predetermined before observing acoustic events. However, the necessary number of acoustic topics for representing acoustic scenes varies in accordance with their contents, and this causes a mismatch between the actual number of classes of acoustic topics and the predetermined number of classes. In our method, the number of classes of acoustic topics can be automatically inferred from sequentially obtained acoustic event sequences on the basis of the online and nonparametric Bayesian technique. The experimental results of online acoustic scene estimation using real-life sounds indicated that the proposed method performed of acoustic scene classification better than the conventional ATM. In addition, the proposed method produced an efficient computation performance.