A Unified Framework for Gesture Recognition and Spatiotemporal Gesture Segmentation

A Unified Framework for Gesture Recognition and Spatiotemporal Gesture Segmentation
复制标题

DOI:
10.1109/tpami.2008.203
复制
发表时间:
2009-09-01
影响因子:
23.6
通讯作者:
Sclaroff, Stan
Sclaroff, Stan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Alon, Jonathan;Athitsos, Vassilis;Sclaroff, Stan

文献摘要

被引文献

相似文献

在手势识别的上下文中,时空手势分割是在视频序列中确定手势手位于何处以及手势何时开始和结束的任务。现有的手势识别方法通常假设已知的空间分割或已知的时间分割,或两者。本文介绍了一个统一的框架,同时进行空间分割,时间分割和识别。在拟议的框架中,信息流动既有自下而上的,也有自上而下的。即使当手的位置高度不明确并且当关于手势何时开始和结束的信息不可用时,也可以识别手势。因此,该方法可以应用于连续图像流,其中在移动的、杂乱的背景前面执行手势。所提出的方法包括三个新的贡献:时空匹配算法,可以容纳多个候选人的手检测在每一帧中,基于分类器的修剪框架,使准确和早期拒绝不良匹配的手势模型,和子手势推理算法,学习哪些手势模型可以错误地匹配其他较长的手势的一部分。该方法的性能进行了评估两个具有挑战性的应用:识别手势的用户穿着短袖衬衫,在前面的一个杂乱的背景,和检索的发生在视频数据库中的兴趣的迹象,包含连续的,未分割的签署美国手语(ASL)的手签名的数字。
Within the context of hand gesture recognition, spatiotemporal gesture segmentation is the task of determining, in a video sequence, where the gesturing hand is located and when the gesture starts and ends. Existing gesture recognition methods typically assume either known spatial segmentation or known temporal segmentation, or both. This paper introduces a unified framework for simultaneously performing spatial segmentation, temporal segmentation, and recognition. In the proposed framework, information flows both bottom-up and top-down. A gesture can be recognized even when the hand location is highly ambiguous and when information about when the gesture begins and ends is unavailable. Thus, the method can be applied to continuous image streams where gestures are performed in front of moving, cluttered backgrounds. The proposed method consists of three novel contributions: a spatiotemporal matching algorithm that can accommodate multiple candidate hand detections in every frame, a classifier-based pruning framework that enables accurate and early rejection of poor matches to gesture models, and a subgesture reasoning algorithm that learns which gesture models can falsely match parts of other longer gestures. The performance of the approach is evaluated on two challenging applications: recognition of hand-signed digits gestured by users wearing short-sleeved shirts, in front of a cluttered background, and retrieval of occurrences of signs of interest in a video database containing continuous, unsegmented signing in American Sign Language (ASL).