Automatic Environmental Sound Recognition: Performance Versus Computational Cost

Automatic Environmental Sound Recognition: Performance Versus Computational Cost
复制标题

DOI:
10.1109/taslp.2016.2592698
复制
发表时间:
2016-07
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Siddharth Sigtia;Adam M. Stark;Sacha Krstulovic;Mark D. Plumbley
Siddharth Sigtia;Adam M. Stark;Sacha Krstulovic;Mark D. Plumbley
中科院分区:
其他
文献类型:
--
作者:
Siddharth Sigtia;Adam M. Stark;Sacha Krstulovic;Mark D. Plumbley

文献摘要

被引文献

相似文献

在物联网的背景下,声音传感应用程序需要在嵌入式平台上运行,在这些平台上,产品定价和外形因素的概念对可用计算能力施加了硬约束。鉴于自动环境声音识别(AESR)算法的发展往往对计算成本的考虑有限,本文通过比较声音分类性能与其计算成本的函数关系来寻找哪种AESR算法能够最大限度地利用有限的计算能力。结果表明,深度神经网络在一定的计算成本范围内提供了最佳的声音分类准确率,而高斯混合模型以始终较小的代价提供了合理的精度,而支持向量机在精度和计算成本之间的折衷方面介于两者之间。
In the context of the Internet of Things, sound sensing applications are required to run on embedded platforms where notions of product pricing and form factor impose hard constraints on the available computing power. Whereas Automatic Environmental Sound Recognition (AESR) algorithms are most often developed with limited consideration for computational cost, this paper seeks which AESR algorithm can make the most of a limited amount of computing power by comparing the sound classification performance as a function of its computational cost. Results suggest that Deep Neural Networks yield the best ratio of sound classification accuracy across a range of computational costs, while Gaussian Mixture Models offer a reasonable accuracy at a consistently small cost, and Support Vector Machines stand between both in terms of compromise between accuracy and computational cost.