Multimodal error correction for speech user interfaces

Multimodal error correction for speech user interfaces
复制标题

DOI:
10.1145/371127.371166
复制
发表时间:
2001-03
期刊:
ACM Trans. Comput. Hum. Interact.
影响因子:
--
通讯作者:
B. Suhm;B. Myers;A. Waibel
B. Suhm;B. Myers;A. Waibel
中科院分区:
其他
文献类型:
--
作者:
B. Suhm;B. Myers;A. Waibel

文献摘要

被引文献

相似文献

尽管商业听写系统和支持语音的电话语音用户界面已经变得容易获得,但语音识别错误仍然是语音用户界面设计和实现中的一个严重问题。先前的研究假设,转换通道可以加快识别错误的交互纠正。本文提出了多模式纠错方法,允许用户在不需要键盘输入的情况下有效地纠正识别错误。通过使用上下文信息来识别校正输入的新型识别算法使校正精度最大化。多模式纠错是在原型多模式听写系统的背景下评估的。研究表明,单峰修正的精确度低于多峰纠错。在听写任务中,多峰校正比单峰校正更快。这项研究还提供了经验证据,表明系统启动的纠错(基于置信度衡量)可能不会加快纠错速度。此外,研究表明,识别的准确性决定了用户在模式之间的选择:虽然用户最初更喜欢语音,但他们学会了通过经验避免无效的校正模式。为了推断这项用户研究的结果,文章引入了一个(基于识别的)多通道交互的性能模型,该模型预测了输入速度,包括纠错所需的时间。将该模型应用于交互式纠错,预测了识别技术的改进对纠错速度的影响,以及识别精度和纠错方法对听写系统生产率的影响。该模型是将多通道交互形式化的第一步。
Although commercial dictation systems and speech-enabled telephone voice user interfaces have become readily available, speech recognition errors remain a serious problem in the design and implementation of speech user interfaces. Previous work hypothesized that switching modality could speed up interactive correction of recognition errors. This article presents multimodal error correction methods that allow the user to correct recognition errors efficiently without keyboard input. Correction accuracy is maximized by novel recognition algorithms that use context information for recognizing correction input. Multimodal error correction is evaluated in the context of a prototype multimodal dictation system. The study shows that unimodal repair is less accurate than multimodal error correction. On a dictation task, multimodal correction is faster than unimodal correction by respeaking. The study also provides empirical evidence that system-initiated error correction (based on confidence measures) may not expedite error correction. Furthermore, the study suggests that recognition accuracy determines user choice between modalities: while users initially prefer speech, they learn to avoid ineffective correction modalities with experience. To extrapolate results from this user study, the article introduces a performance model of (recognition-based) multimodal interaction that predicts input speed including time needed for error correction. Applied to interactive error correction, the model predicts the impact of improvements in recognition technology on correction speeds, and the influence of recognition accuracy and correction method on the productivity of dictation systems. This model is a first step toward formalizing multimodal interaction.