ITR [ASE+ECS] [dmc+int]: COLLABORATIVE RESEARCH: DDDAS Advances in Recognition and Interpretation of Human Motion: An Integrated Approach to ASL Recognition
ITR [ASE+ECS] [dmc+int]: COLLABORATIVE RESEARCH: DDDAS Advances in Recognition and Interpretation of Human Motion: An Integrated Approach to ASL Recognition
批准号:
0427267
负责人:
Christian Vogler
金额:
$25.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-10-15 至 2007-09-30
中文摘要
NSF ITR标题:ITR -[ASE+ECS] -[dmc+int]:DDDAS在人体运动识别和解释方面的进展:ASL识别的综合方法该项目旨在推进基于计算机的美国手语(ASL)识别领域的最新技术。迄今为止,手语识别主要集中在检测主要用手臂和手表达的单个符号(单词)。这是一个主要的限制,因为关键的语言信息——包括语法特征,如否定、一致和问题状态——是通过“非手动”语言标记传达的。这些非手动标记包括面部表情(如扬起或下垂的眉毛、不同的凝视和眼睛的光圈、鼻子的皱纹和嘴巴的运动)和手势或头部的周期性运动(如倾斜、点头和摇晃)。如果没有对人工和非人工产生的语言信息进行适当的建模,任何手语识别或生成系统都不可能成功。事实上,这些关键的非手动行为与手动签名并行发生,并且它们暂时与短语而不是单个签名对齐,这极大地使任务复杂化。进一步的问题出现是因为很难从视频中跟踪人类面部运动的微小细节,以及不同个体对手势和非手动语言标记的具体实现(风格)的差异,就像个体产生给定口语的具体方式存在差异一样。因此,一种全面的美国手语识别方法需要整合来自不同空间和时间尺度的多个数据源的信息,应用关于美国手语的手动和非手动方面的语言知识,并对手动和非手动渠道中活动的相互依赖性进行建模。这个合作项目汇集了计算机视觉、语言学和识别领域的研究人员的专业知识,以实现其目标。在计算机视觉方面,主要研究人员将研究使用局部自由变形和新颖的配准方法来增强我们现有的面部跟踪软件,从而捕捉面部运动的微小细节,并提高跟踪的鲁棒性。跟踪过程中会产生大量的人脸参数,研究人员提出通过非线性子空间流形嵌入来减少人脸参数。这种嵌入降低了参数空间的维数,更重要的是,还导致了样式和内容的分离。虽然样式特定于每个签名者,但内容捕获所有签名者的共性。因此,通过关注内容组件,pi期望能够克服不同签名者之间的差异,并执行独立于签名者的识别。在识别方面,研究人员将把面部微动作的语言学知识与计算聚类方法结合起来,开发识别所需的统计模型。最初,这些将基于本研究小组和其他地方先前工作的隐马尔可夫模型;然而,它们描述人类运动动力学方面的能力是有限的。为了克服这些限制,pi将研究切换线性动态系统的使用,辅以耦合动态贝叶斯网络来建模和捕获同时发生的微动作的相互作用。语言学家和计算机科学家将合作探索利用美国手语语言组织信息来改进识别策略的最佳方法。这项研究将在国家手语和手势资源中心现有的语言注释语料库上进行,以及从5-8名美国手语母语者中收集的新数据,这些数据也将在项目过程中进行注释。注释将用于语言建模;它们还为执行和验证计算机视觉和识别研究提供了“基本事实”。更广泛的影响:基于计算机的美国手语技术可以扩展到更通用的系统,用于手语识别和生成,以及其他类型的人类动作的解释,例如用于HCI的面部手势识别,监视,身份验证,审讯,访谈和医疗诊断应用。分发的材料将使语言学、计算机科学和其他领域的研究人员受益。对聋人进行小学和中学教育以及培训手语翻译人员有直接的需求。多媒体(语言)信息技术的改进有望为聋人提供更多的就业机会,并改善他们接受职业和高等教育的机会。最后,这个项目本身将在教育、意识和鼓励聋哑学生方面提供巨大的推动,使他们能够从事直接影响他们和他们的社区的前沿研究。
英文摘要
NSF ITR title: ITR -[ASE+ECS] - [dmc+int]:DDDAS Advances in Recognition and Interpretation of Human Motion:An Integrated Approach to ASL RecognitionThis project is aimed at advancing the state of the art in the field of computer-based American Sign Language (ASL) recognition. To date, sign language recognition has focused primarily on detecting individual signs (words), articulated primarily with the arms and hands. This is a major limitation, given that critical linguistic information-including grammatical features such as negation, agreement, and question status-is conveyed through "non-manual" linguistic markings. These non-manual markings include facial expressions (such as raised or lowered eyebrows, varying gaze and aperture of the eyes, wrinkling of the nose, and mouth movements) and gestures or periodic movements of the head (such as tilts, nods, and shakes). No system for sign language recognition or generation can succeed without properly modeling the linguistic information produced both manually and non-manually.The fact that these critical non-manual behaviors occur in parallel with manual signing, and that they are temporally aligned with phrases rather than with individual signs, greatly complicates the task. Further problems arise because of the difficulties of tracking the minute details of human facial movements from video, and the variations in the specific realizations (style) of manual signs and non-manual linguistic markings across different individuals, just as there is variation in the specific ways in which individuals produce a given spoken language. A comprehensive approach to ASL recognition thus requires the integration of information from multiple data sources with different spatial and temporal scales, the application of linguistic knowledge about both the manual and the non-manual aspects of ASL, and the modeling of interdependencies of activities in the manual and non-manual channels.This collaborative project brings together the expertise of researchers in the fields of computer vision, linguistics, and recognition to achieve its goals. On the computer vision side, the principal investigators (PIs) will investigate the use of local free-form deformations and novel registration methods to enhance our existing face tracking software, so as to capture the minute details of the facial movements, and to improve robustness of the tracking. The tracking process results in a large number of facial parameters, which the researchers propose to reduce through nonlinear subspace manifold embedding. This embedding reduces the dimensionality of the parameter space, and more importantly, also results in a separation of style and content. Whereas style is specific to each signer, content captures the commonalities across all signers. Hence, by focusing on the content component, the PIs expect to be able to overcome the variations across signers and perform signer-independent recognition.On the recognition side, the researhcers will combine the linguistic knowledge about facial microactions with computational clustering approaches to develop the necessary statistical models for recognition. Initially, these will be based on Hidden Markov Models from previous work by this research group and elsewhere; however, their power to describe the dynamical aspects of human movements is limited. To overcome these limitations, the PIs will research the use of Switching Linear Dynamic Systems, augmented by Coupled Dynamic Bayesian Networks to model and capture the interactions of the simultaneously occurring microactions. Linguists and computer scientists will collaborate in exploring the best ways to leverage information about the linguistic organization of ASL for improvement of recognition strategies.This research will be performed on the existing linguistically annotated corpus of the National Center for Sign Language and Gesture Resources, as well as new data to be collected from 5-8 native ASL signers, which will also be annotated over the course of the project. The annotations will be used for the linguistic modeling; they also provide the "ground truth" for performing and validating the computer vision and recognition research.Broader impact: The computer-based techniques for ASL can be extended to more general systems for sign language recognition and generation, as well as for interpretation of other types of human movements, such as face gesture recognition for HCI, surveillance, verification of identity, interrogation, interviews and medical diagnosis applications. The materials to be distributed will benefit researchers in linguistics, computer science, and other domains. There are immediate applications for primary and secondary education of the deaf and training of sign language interpreters. Improvements in multimedia (linguistic) information technology promise to offer expanded employment possibilities for the deaf, as well as improved access to vocational and post-secondary education. Finally, this project itself will provide a huge boost in terms of education, awareness and encouragement of deaf students in enabling them to work on cutting-edge research that directly affects them and their community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金