Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition

Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Dimitris N. Metaxas;Mark Dilsizian;C. Neidle
Dimitris N. Metaxas;Mark Dilsizian;C. Neidle
中科院分区:
其他
文献类型:
--
作者:
Dimitris N. Metaxas;Mark Dilsizian;C. Neidle

文献摘要

相似文献

我们提出了一种新的通用框架,用于使用有限数量的标注数据从单目视频中识别手势。我们在这里描述的混合框架的新奇之处在于,我们利用了最先进的学习方法,同时也融入了基于我们所知的词汇符号的语言组成的特征。特别是,我们分析手形、方向、位置和运动轨迹,然后使用CRF来组合这些具有语言意义的信息以用于手势识别。与纯粹的数据驱动方法相比,我们对符号生成的这些子组件的健壮建模和识别允许对符号识别问题进行有效的参数化。这种参数化实现了一种可伸缩和可扩展的时间序列学习方法,该方法推进了手势识别的技术水平,如本文报告的对美国手语(ASL)中孤立的、引文形式的词汇手势的识别结果所示。
We introduce a new general framework for sign recognition from monocular video using limited quantities of annotated data. The novelty of the hybrid framework we describe here is that we exploit state-of-the art learning methods while also incorporating features based on what we know about the linguistic composition of lexical signs. In particular, we analyze hand shape, orientation, location, and motion trajectories, and then use CRFs to combine this linguistically significant information for purposes of sign recognition. Our robust modeling and recognition of these sub-components of sign production allow an efficient parameterization of the sign recognition problem as compared with purely data-driven methods. This parameterization enables a scalable and extendable time-series learning approach that advances the state of the art in sign recognition, as shown by the results reported here for recognition of isolated, citation-form, lexical signs from American Sign Language (ASL).