CHS: Medium: Collaborative Research: Scalable Integration of Data-Driven and Model-Based Methods for Large Vocabulary Sign Recognition and Search
CHS: Medium: Collaborative Research: Scalable Integration of Data-Driven and Model-Based Methods for Large Vocabulary Sign Recognition and Search
批准号:
1763569
负责人:
Matt Huenerfauth
金额:
$20.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2023-07-31
中文摘要
在美国手语(ASL)中查找一个不熟悉的手势出人意料地困难。大多数美国手语词典按照字母顺序列出了大致的英文翻译,因此不理解或不知道其英文翻译的用户将不知道如何找到它。ASL缺乏书面形式或基于这样的书写系统的直观的“字母排序”。尽管一些词典提供了搜索标志的替代方法,但基于各种属性的明确规范,用户仍然经常必须查看数百张标志的图片以找到与不熟悉的标志的匹配(如果它在该词典中存在的话)。这项研究将创建一个框架,能够开发一个用户友好的、基于视频的手语查找界面,用于在线美国手语视频词典和资源,并为美国手语注释提供便利。输入将包括用户对标志的网络摄像头记录,或用户从数字视频中识别标志的开始和结束帧。为了测试新工具在现实应用中的有效性,该团队将与领先的高中和大学ASL教学材料生产商合作,该公司正在开发第一个基于视频的ASL手语解释的多媒体ASL词典。该查找界面将用于在波士顿大学和RIT的ASL班级中实验性地搜索ASL词典。项目成果将彻底改变聋儿、学习ASL的学生或有聋儿的家庭查找ASL词典的方式。他们将通过基于视频的标志查找提高ASL视频注释的效率、准确性和一致性,从而加快ASL语言学和技术的研究。他们还将为未来造福聋人用户的技术奠定基础,例如通过ASL视频集合的视频搜索,或ASL到英语的翻译,手势识别是其先驱。新的带有语言注释的视频数据和软件工具将被公开分享,供其他人在语言和计算机科学研究以及教育中使用。由于从2D视频中识别3D结构所涉及的非线性,以及手语的复杂语言组织,从视频中识别手势仍然是一个开放而困难的问题。与手势产生和辨别相关的语言参数包括手形和方位,相对于身体或手势空间的位置,运动轨迹,在某些情况下,面部表情/头部运动。另一个复杂的问题是,属于不同类别的符号具有不同的内部结构,因此受到不同的语言限制,需要不同的识别策略;然而,以往的研究通常未能解决这些差异。这些挑战由于签名者之间和签名者内部的差异,以及在连续签名中,关于上述几个参数的共同发音效应(即,来自相邻手势的影响)而变得更加复杂。纯数据驱动的方法不适合手势识别,因为可用的、一致注释的数据数量有限,而且涉及的语言结构复杂,很难推断。因此,以前的研究通常集中在问题的选定方面,往往将工作限制在有限的词汇表中,因此导致方法不可扩展。更重要的是,很少有方法涉及4D(时空)建模和对特定类型符号的语言属性的关注。需要一种基于计算机从视频中识别ASL的新方法。在这项研究中,方法将是建立一个新的混合的,可扩展的,用于从大量词汇中识别手势的计算框架,这是以前从未实现的。这项研究将战略性地结合最先进的计算机视觉、机器学习方法和语言建模。它将利用团队现有的公开共享的ASL语料库和Sign Bank-由本地签名者制作的经过语言注释和分类的视频记录-将进行扩充,以满足本项目的要求。该奖项反映了NSF的法定使命,并已通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
It is surprisingly difficult to look up an unfamiliar sign in American Sign Language (ASL). Most ASL dictionaries list signs in alphabetical order based on approximate English translations, so a user who does not understand a sign or know its English translation would not know how to find it. ASL lacks a written form or intuitive "alphabetical sorting" based on such a writing system. Although some dictionaries make available alternative ways to search for a sign, based on explicit specification of various properties, a user must often still look through hundreds of pictures of signs to find a match to the unfamiliar sign (if it is present at all in that dictionary). This research will create a framework that will enable the development of a user-friendly, video-based sign-lookup interface, for use with online ASL video dictionaries and resources, and for facilitation of ASL annotation. Input will consist of either a webcam recording of a sign by the user, or user identification of the start and end frames of a sign from a digital video. To test the efficacy of the new tools in real-world applications, the team will partner with the leading producer of pedagogical materials for ASL instruction in high schools and colleges, which is developing the first multimedia ASL dictionary with video-based ASL definitions for signs. The lookup interface will be used experimentally to search the ASL dictionary in ASL classes at Boston University and RIT. Project outcomes will revolutionize how deaf children, students learning ASL, or families with deaf children search ASL dictionaries. They will accelerate research on ASL linguistics and technology, by increasing efficiency, accuracy, and consistency of annotations of ASL videos through video-based sign lookup. And they will lay the groundwork for future technologies to benefit deaf users, such as search by video example through ASL video collections, or ASL-to-English translation, for which sign-recognition is a precursor. The new linguistically annotated video data and software tools will be shared publicly, for use by others in linguistic and computer science research, as well as in education. Sign recognition from video is still an open and difficult problem because of the nonlinearities involved in recognizing 3D structures from 2D video, and the complex linguistic organization of sign languages. The linguistic parameters relevant to sign production and discrimination include hand configuration and orientation, location relative to the body or in signing space, movement trajectory, and in some cases, facial expressions/head movements. An additional complication is that signs belonging to different classes have distinct internal structures, and are thus subject to different linguistic constraints and require distinct recognition strategies; yet prior research has generally failed to address these distinctions. The challenges are compounded by inter- and intra- signer variations, and, in continuous signing, by co-articulation effects (i.e., influence from adjacent signs) with respect to several of the above parameters. Purely data-driven approaches are ill-suited to sign recognition given the limited quantities of available, consistently annotated data and the complexity of the linguistic structures involved, which are hard to infer. Prior research has, for this reason, generally focused on selected aspects of the problem, often restricting the work to a limited vocabulary, and therefore resulting in methods that are not scalable. More importantly, few if any methods involve 4D (spatio-temporal) modeling and attention to the linguistic properties of specific types of signs. A new approach to computer-based recognition of ASL from video is needed. In this research, the approach will be to build a new hybrid, scalable, computational framework for sign identification from a large vocabulary, which has never before been achieved. This research will strategically combine state-of-the-art computer vision, machine-learning methods, and linguistic modeling. It will leverage the team's existing publicly shared ASL corpora and Sign Bank - linguistically annotated and categorized video recordings produced by native signers - which will be augmented to meet the requirements of this project.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Interest and Requirements for Sound-Awareness Technologies Among Deaf and Hard-of-Hearing Users of Assistive Listening Devices
助听设备聋人和听力障碍用户对声音感知技术的兴趣和要求
DOI:
10.1007/978-3-030-49108-6_11
发表时间:
2020
期刊:
Universal Access in Human-Computer Interaction. Applications and Practice. HCII 2020. Lecture Notes in Computer Science
影响因子:
--
作者:
[Yeung, Peter, Alonzo, Oliver, Huenerfauth, Matt]
通讯作者:
Huenerfauth, Matt
Support in the Moment: Benefits and use of video-span selection and search for sign-language video comprehension among ASL learners
当前支持:视频范围选择和搜索对 ASL 学习者手语视频理解的好处和使用
DOI:
10.1145/3517428.3544883
发表时间:
2022
期刊:
Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility
影响因子:
--
作者:
[Hassan, Saad, Amin, Akhter Al, de Lacerda Pataca, Caluã, Navarro, Diego, Gordon, Alexis, Lee, Sooyeon, Huenerfauth, Matt]
通讯作者:
Huenerfauth, Matt
DOI:
10.1145/3491102.3501986
发表时间:
2022-04
期刊:
Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems
影响因子:
--
作者:
[Saad Hassan;Akhter Al Amin;Alexis Gordon;Sooyeon Lee;Matt Huenerfauth]
通讯作者:
Saad Hassan;Akhter Al Amin;Alexis Gordon;Sooyeon Lee;Matt Huenerfauth
Understanding ASL Learners’ Preferences for a Sign Language Recording and Automatic Feedback System to Support Self-Study
了解 ASL 学习者对支持自学的手语录音和自动反馈系统的偏好
DOI:
10.1145/3517428.3550367
发表时间:
2022
期刊:
Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility
影响因子:
--
作者:
[Hassan, Saad, Lee, Sooyeon, Metaxas, Dimitris, Neidle, Carol, Huenerfauth, Matt]
通讯作者:
Huenerfauth, Matt
Designing and Experimentally Evaluating a Video-based American Sign Language Look-up System
基于视频的美国手语查找系统的设计和实验评估
DOI:
10.1145/3498366.3505804
发表时间:
2022
期刊:
ACM SIGIR Conference on Human Information Interaction and Retrieval
影响因子:
--
作者:
[Hassan, Saad]
通讯作者:
Hassan, Saad
共 9 条
Collaborative Research: HCC: Medium: Linguistically-Driven Sign Recognition from Continuous Signing for American Sign Language (ASL)
-
批准号:2212303
-
项目类别:Standard Grant
-
资助金额:$16.5万
-
财政年份:2022
-
负责人:Matt Huenerfauth
-
依托单位:
CHS: Medium: Critical Factors for Automatic Speech Recognition in Supporting Small Group Communication Between People who are Deaf or Hard of Hearing and Hearing Colleagues
-
批准号:1954284
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2020
-
负责人:Matt Huenerfauth
-
依托单位:
Collaborative Research: Automatic Text-Simplification and Reading-Assistance to Support Self-Directed Learning by Deaf and Hard-of-Hearing Computing Workers
-
批准号:1822747
-
项目类别:Standard Grant
-
资助金额:$39.19万
-
财政年份:2018
-
负责人:Matt Huenerfauth
-
依托单位:
CRII: CHS: Augmented Fabrication for Non-Expert Users of Digital Fabrication Systems
-
批准号:1464377
-
项目类别:Continuing Grant
-
资助金额:$17.5万
-
财政年份:2015
-
负责人:Matt Huenerfauth
-
依托单位:
CCE STEM: Ethical Inclusion of People with Disabilities through Undergraduate Computing Education
-
批准号:1540396
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2015
-
负责人:Matt Huenerfauth
-
依托单位:
CHS: Medium: Collaborative Research: Immediate Feedback to Support Learning American Sign Language through Multisensory Recognition
-
批准号:1400906
-
项目类别:Standard Grant
-
资助金额:$53.8万
-
财政年份:2014
-
负责人:Matt Huenerfauth
-
依托单位:
CHS: Medium: Collaborative Research: Immediate Feedback to Support Learning American Sign Language through Multisensory Recognition
-
批准号:1462280
-
项目类别:Standard Grant
-
资助金额:$53.8万
-
财政年份:2014
-
负责人:Matt Huenerfauth
-
依托单位:
HCC: Medium: Collaborative Research: Generating Accurate, Understandable Sign Language Animations Based on Analysis of Human Signing
-
批准号:1506786
-
项目类别:Continuing Grant
-
资助金额:$6.0万
-
财政年份:2014
-
负责人:Matt Huenerfauth
-
依托单位:
HCC: Medium: Collaborative Research: Generating Accurate, Understandable Sign Language Animations Based on Analysis of Human Signing
-
批准号:1065009
-
项目类别:Continuing Grant
-
资助金额:$23.22万
-
财政年份:2011
-
负责人:Matt Huenerfauth
-
依托单位:
Doctoral Consortium for ASSETS 2010
-
批准号:1035382
-
项目类别:Standard Grant
-
资助金额:$2.72万
-
财政年份:2010
-
负责人:Matt Huenerfauth
-
依托单位:
CAREER: Learning to Generate American Sign Language Animation through Motion-Capture and Participation of Native ASL Signers
-
批准号:0746556
-
项目类别:Continuing Grant
-
资助金额:$58.15万
-
财政年份:2008
-
负责人:Matt Huenerfauth
-
依托单位:
海外基金