Toward Natual Gesture/Speech HCI: A Case Study of Weather Narration

Toward Natual Gesture/Speech HCI: A Case Study of Weather Narration
复制标题

迈向自然手势/语音人机交互:天气叙事案例研究

DOI:
--
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
Yogesh Sethi
Yogesh Sethi
中科院分区:
--
文献类型:
--
作者:
E. Ozyildiz;Indrajit Poddar;Rajeev Sharma;Yogesh Sethi

文献摘要

被引文献

相似文献

为了将自然感融入到人机界面(HCI)的设计中,需要开发能够处理连续自然手势和语音输入的识别技术。隐马尔可夫模型(HMM)为连续手势识别和多模式融合提供了一个很好的框架[11]。许多不同的研究人员[13,12,2]报告了使用HMM进行手势识别的高识别率[9]。然而,他们用于识别的手势是精确定义的,并受到句法和语法限制。但自然的手势在句法上不会连在一起[3]。此外,对自然手势进行严格的分类是不可行的。在这篇文章中,我们研究了在一个非常自然的领域中制作的手势,即天气人在天气地图前讲述的手势。天气预报员的手势被嵌入到解说词中。这为我们提供了来自非受控环境的丰富数据,以研究显示环境中语音和手势之间的相互作用。我们假设这个域非常类似于一个自然的人机界面。我们实现了一个基于连续隐马尔可夫模型的手势识别框架。为了了解手势和语音之间的相互作用,我们用一些口语关键词对不同的手势进行了共现分析。我们还展示了基于共现分析改进连续手势识别结果的可能性。通过对彩色分割的视频图像流使用预测卡尔曼滤波来实现快速的特征提取和跟踪。天气领域的结果应该是朝着自然手势/语音人机交互迈出的一步。
In order to incorporate naturalness in the design of Human Computer Interfaces (HCI), it is desirable to develop recognition techniques capable of handling continuous natural gesture and speech inputs. Hidden Markov Models (HMMs) provide a good framework for continuous gesture recognition and also for multimodal fusion [11]. Many different researchers [13, 12, 2], have reported high recognition rates for gesture recognition using HMMs [9]. However the gestures which were used for recognition by them were defined precisely and were bound with syntactical and grammatical constraints. But natural gestures do not string together in syntactical bindings [3]. Moreover strict classification of natural gesture is not feasible. In this paper we have examined hand gestures made in a very natural domain, that of a weather person narrating in front of a weather map. The gestures made by the weather person are embedded in a narration. This provides us with abundant data from an uncontrolled environment to study the interaction between speech and gesture in the context of a display. We hypothesize that this domain is very similar to that of a natural HCI interface. We have implemented a continuous HMM based gesture recognition framework. In order to understand the interaction between the gesture and speech, we have done a co-occurrence analysis of different gestures with some spoken keywords. We have also shown the possibility of improving continuous gesture recognition results based on the co-occurrence analysis. Fast feature extraction and tracking is accomplished by the use of a predictive Kalman filtering on color segmented stream of video images. The results in the weather domain should be a step toward natural gesture/speech HCI.