Evaluating the impact of pushing voice-driven interaction pipelines to the edge

Evaluating the impact of pushing voice-driven interaction pipelines to the edge
复制标题

DOI:
10.1145/3203217.3203242
复制
发表时间:
2018-05
期刊:
Proceedings of the 15th ACM International Conference on Computing Frontiers
影响因子:
--
通讯作者:
S. Sridhar;Matthew E. Tolentino
S. Sridhar;Matthew E. Tolentino
中科院分区:
其他
文献类型:
--
作者:
S. Sridhar;Matthew E. Tolentino

文献摘要

相似文献

随着Alexa Voice Services和Google Home的发布,语音驱动的交互式计算已经迅速变得司空见惯。语音交互应用程序包含多个组件,包括复杂的语音识别和翻译算法,自然语言理解和生成功能,以及通常称为技能的自定义计算功能。语音驱动的交互式系统由使用这些组件的软件管道组成。这些管道通常是资源密集型的,必须快速执行以保持对话一致的延迟;因此,语音交互管道通常在云中计算。然而,在许多情况下,云连接可能并不实用,因此需要在边缘执行这些语音交互管道。在本文中,我们评估了将语音交互管道推到资源受限的边缘设备的可行性。在紧急情况下,当连接到云是不切实际的时候,为第一响应者启用语音驱动的接口的目标驱动下,我们描述了一个完整的开源语音交互管道的端到端性能,用于四种不同的配置,从完全基于云到完全基于边缘。然后,我们确定并评估了几种优化,例如缓存和定制的声学模型,这些模型使语音驱动的交互管道能够在计算能力较弱的边缘设备上以比使用高性能云资源更低的响应延迟完全执行。
With the releases of Alexa Voice Services and Google Home, voice-driven interactive computing has quickly become commonplace. Voice interactive applications incorporate multiple components including complex speech recognition and translation algorithms, natural language understanding and generation capabilities, as well as custom compute functions commonly referred to as skills. Voice-driven interactive systems are composed of software pipelines using these components. These pipelines are typically resource intensive and must be executed quickly to maintain dialogue-consistent latency; consequently, voice interaction pipelines are usually computed in the cloud. However, for many cases, cloud connectivity may not be practical and thus require these voice interactive pipelines be executed at the edge. In this paper, we evaluate the feasibility of pushing voice interaction pipelines to resource constrained edge devices. Driven by the goal of enabling voice-driven interfaces for first responders during emergencies when connectivity to the cloud is impractical, we characterize the end-to-end performance of a complete open source voice interaction pipeline for four different configurations ranging from entirely cloud-based to completely edge-based. We then identify and evaluate several optimizations, such as caching and customized acoustic models that enable voice-driven interaction pipelines to be fully executed at computationally-weak edge devices at lower response latencies than using high-performance cloud resources.