Automatic speech recognition and speech variability: A review

Automatic speech recognition and speech variability: A review
复制标题

DOI:
10.1016/j.specom.2007.02.006
复制
发表时间:
2007-10-01
影响因子:
3.2
通讯作者:
Wellekens, C.
Wellekens, C.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Benzeghiba, M.;De Mori, R.;Wellekens, C.

文献摘要

被引文献

相似文献

在自动语音识别(ASR)和口语系统的技术和开发方面,经常取得重大进展。然而,在某些情况下,灵活的解决方案和用户满意度仍然存在技术障碍。这与语音对环境(背景噪声)的敏感性、语法和语义知识的弱表达等因素有关。目前的研究也强调了在处理语音中自然存在的变异方面的不足。例如,对外国口音缺乏鲁棒性,排除了特定人群的使用。此外,一些应用,如目录辅助,特别强调核心识别技术,由于非常高的活动词汇(应用困惑)。影响言语实现的因素有很多:地域因素、社会语言因素、环境因素、说话人自身因素等。这些产生了可能无法正确建模的各种变化(说话者、性别、语速、发声努力、地区口音、说话风格、非平稳性等),特别是在系统培训资源稀缺的情况下。本文概述了目前与这些主题有关的进展。(C)2007 Elsevier B.V.保留所有权利。
Major progress is being recorded regularly on both the technology and exploitation of automatic speech recognition (ASR) and spoken language systems. However, there are still technological barriers to flexible solutions and user satisfaction under some circumstances. This is related to several factors, such as the sensitivity to the environment (background noise), or the weak representation of grammatical and semantic knowledge.Current research is also emphasizing deficiencies in dealing with variation naturally present in speech. For instance, the lack of robustness to foreign accents precludes the use by specific populations. Also, some applications, like directory assistance, particularly stress the core recognition technology due to the very high active vocabulary (application perplexity). There are actually many factors affecting the speech realization: regional, sociolinguistic, or related to the environment or the speaker herself. These create a wide range of variations that may not be modeled correctly (speaker, gender, speaking rate, vocal effort, regional accent, speaking style, non-stationarity, etc.), especially when resources for system training are scarce. This paper outlines current advances related to these topics. (C) 2007 Elsevier B.V. All rights reserved.