Evaluating Voice Interaction Pipelines at the Edge

Evaluating Voice Interaction Pipelines at the Edge
复制标题

DOI:
10.1109/ieee.edge.2017.46
复制
发表时间:
2017-06
期刊:
2017 IEEE International Conference on Edge Computing (EDGE)
影响因子:
--
通讯作者:
S. Sridhar;Matthew E. Tolentino
S. Sridhar;Matthew E. Tolentino
中科院分区:
其他
文献类型:
--
作者:
S. Sridhar;Matthew E. Tolentino

文献摘要

相似文献

随着Alexa语音服务和Google Home的发布,语音驱动的交互式计算很快变得司空见惯。语音交互应用程序包含多个组件,包括复杂的语音识别和翻译算法,自然语言理解和生成功能,以及通常称为技能的自定义计算功能。语音驱动的交互式系统由使用这些组件的软件管道组成。这些管道通常是资源密集型的,并且必须快速执行以维护对话一致的延迟。因此,语音交互管道通常完全在云中计算。然而,在许多情况下,云连接可能并不实用,并且需要在边缘执行这些语音交互管道。在本文中,我们评估了将语音驱动的管道推向计算薄弱边缘设备的影响。我们的主要动机是在紧急情况下为第一响应者启用语音驱动的界面,例如建筑物火灾,当连接到云是不切实际的。我们首先描述了一个完整的开源语音交互管道的端到端性能,用于四种不同的配置,从完全基于云到完全基于边缘。我们还确定了潜在的优化机会,使语音驱动交互管道能够在计算能力较弱的边缘设备上以比高性能云服务更低的响应延迟完全执行
With the recent releases of Alexa Voice Services and Google Home, voice-driven interactive computing is quickly become commonplace. Voice interactive applications incorporate multiple components including complex speech recognition and translation algorithms, natural language understanding and generation capabilities, as well as custom compute functions commonly referred to as skills. Voice-driven interactive systems are composed of software pipelines using these components. These pipelines are typically resource intensive and must be executed quickly to maintain dialogue-consistent latencies. Consequently, voice interaction pipelines are usually computed entirely in the cloud. However, for many cases, cloud connectivity may not be practical and require these voice interactive pipelines be executed at the edge. In this paper, we evaluate the impact of pushing voice-driven pipelines to computationally-weak edge devices. Our primary motivation is to enable voice-driven interfaces for first responders during emergencies, such as building fires, when connectivity to the cloud is impractical. We first characterize the end-to-end performance of a complete open source voice interaction pipeline for four different configurations ranging from entirely cloud-based to completely edge-based. We also identify potential optimization opportunities to enable voice-drive interaction pipelines to be fully executed at computationally-weak edge devices at lower response latencies than high-performance cloud services