Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification

Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification
复制标题

DOI:
10.1109/icassp39728.2021.9414287
复制
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Sangeeta Srivastava;Dhrubojyoti Roy;M. Cartwright;J. Bello;A. Arora
Sangeeta Srivastava;Dhrubojyoti Roy;M. Cartwright;J. Bello;A. Arora
中科院分区:
其他
文献类型:
--
作者:
Sangeeta Srivastava;Dhrubojyoti Roy;M. Cartwright;J. Bello;A. Arora

文献摘要

相似文献

将语义信息编码到低维向量表示中的嵌入模型在具有有限训练数据的各种机器学习任务中是有用的。然而,这些模型通常太大,无法支持小型边缘设备中的推理,这促使通过知识蒸馏(KD)来训练较小但预测性较强的学生嵌入模型。虽然知识蒸馏传统上使用教师的原始训练数据集来训练学生,但我们假设使用与学生的目标域类似的数据集可以更好地压缩和训练所述域的效率,但代价是降低了其他(非相关)域的通用性。因此,我们引入了专用嵌入近似(SEA)来训练学生特征化器来近似给定目标域的教师嵌入流形。我们证明了SEA在城市噪声监测的声学事件分类背景下的可行性,并表明利用与此目标域相关的数据集不仅提高了原始嵌入模型的基线性能,而且还产生了具有竞争力的学生,其存储和激活记忆的数量级减少了>1个数量级。我们进一步研究使用随机和知情的抽样技术的影响,在SEA降维。
Embedding models that encode semantic information into low-dimensional vector representations are useful in various machine learning tasks with limited training data. However, these models are typically too large to support inference in small edge devices, which motivates training of smaller yet comparably predictive student embedding models through knowledge distillation (KD). While knowledge distillation traditionally uses the teacher’s original training dataset to train the student, we hypothesize that using a dataset similar to the student’s target domain allows for better compression and training efficiency for the said domain, at the cost of reduced generality across other (non-pertinent) domains. Hence, we introduce Specialized Embedding Approximation (SEA) to train a student featurizer to approximate the teacher’s embedding manifold for a given target domain. We demonstrate the feasibility of SEA in the context of acoustic event classification for urban noise monitoring and show that leveraging a dataset related to this target domain not only improves the baseline performance of the original embedding model but also yields competitive students with >1 order of magnitude lesser storage and activation memory. We further investigate the impact of using random and informed sampling techniques for dimensionality reduction in SEA.