ERSAM: Neural Architecture Search for Energy-Efficient and Real-Time Social Ambiance Measurement

ERSAM: Neural Architecture Search for Energy-Efficient and Real-Time Social Ambiance Measurement
复制标题

DOI:
10.1109/icassp49357.2023.10095360
复制
发表时间:
2023-03
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Chaojian Li;Wenwan Chen;Jiayi Yuan;Yingyan Lin;Ashutosh Sabharwal
Chaojian Li;Wenwan Chen;Jiayi Yuan;Yingyan Lin;Ashutosh Sabharwal
中科院分区:
其他
文献类型:
--
作者:
Chaojian Li;Wenwan Chen;Jiayi Yuan;Yingyan Lin;Ashutosh Sabharwal

文献摘要

相似文献

社交氛围描述了社交互动发生的背景,并且可以通过计算同时发言者的数量使用语音音频来测量。这种测量已经实现了各种心理健康跟踪和以人为中心的物联网应用。虽然设备上的社交氛围测量(SAM)对于确保用户隐私并因此促进上述应用的广泛采用是非常期望的,但是最先进的深度神经网络(DNN)驱动的SAM解决方案所需的计算复杂度与移动的设备上通常受限的资源不一致。此外,由于各种隐私限制和所需的人工努力,在临床环境下只有有限的标记数据可用或实用,进一步挑战了设备上SAM解决方案的可实现精度。为此,我们提出了一个专用的神经架构搜索框架的节能和实时SAM(ERSAM)。具体来说,我们的ERSAM框架可以自动搜索DNN,从而推动移动的SAM解决方案的可实现精度与硬件效率的边界。例如,ERSAM交付的DNN在Pixel 3手机上仅消耗40 mW · 12 h能量和0.05秒处理延迟,而在LibriSpeech生成的社交氛围数据集上仅实现14.3%的错误率。我们可以预期,我们的ERSAM框架可以为无处不在的设备上SAM解决方案铺平道路,这些解决方案的需求不断增长。
Social ambiance describes the context in which social interactions happen, and can be measured using speech audio by counting the number of concurrent speakers. This measurement has enabled various mental health tracking and human-centric IoT applications. While on-device Socal Ambiance Measure (SAM) is highly desirable to ensure user privacy and thus facilitate wide adoption of the aforementioned applications, the required computational complexity of state-of-the-art deep neural networks (DNNs) powered SAM solutions stands at odds with the often constrained resources on mobile devices. Furthermore, only limited labeled data is available or practical when it comes to SAM under clinical settings due to various privacy constraints and the required human effort, further challenging the achievable accuracy of on-device SAM solutions. To this end, we propose a dedicated neural architecture search framework for Energy-efficient and Real-time SAM (ERSAM). Specifically, our ERSAM framework can automatically search for DNNs that push forward the achievable accuracy vs. hardware efficiency frontier of mobile SAM solutions. For example, ERSAM-delivered DNNs only consume 40 mW • 12 h energy and 0.05 seconds processing latency for a 5 seconds audio segment on a Pixel 3 phone, while only achieving an error rate of 14.3% on a social ambiance dataset generated by LibriSpeech. We can expect that our ERSAM framework can pave the way for ubiquitous on-device SAM solutions which are in growing demand.