课题基金 / 基金详情

HCC: Medium: Collaborative Research: Generating Accurate, Understandable Sign Language Animations Based on Analysis of Human Signing

HCC: Medium: Collaborative Research: Generating Accurate, Understandable Sign Language Animations Based on Analysis of Human Signing
HCC:媒介:协作研究:根据人类手语分析生成准确、可理解的手语动画
批准号:
1064965
负责人:
Dimitris Metaxas
金额:
$47.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-07-01 至 2016-06-30

项目摘要

项目成果

Dimitris Metaxas的其他基金

相似基金

相关文献

中文摘要
翻译
美国手语(ASL)动画有可能使美国许多只具有有限英语读写能力的聋人成年人获得信息。在这项涉及三个机构合作的研究中,pi的目标是通过计算技术更好地理解美国手语语言学,同时推进为聋人无障碍应用程序生成美国手语动画的最新技术。为了达到这些目的,pi将开发基于语言的ASL生成的两个方面的模型:头部手势和面部表情所需的动作,这些动作携带重要的语法信息,并且经常扩展到比单个手势更大的领域,以及ASL签名的手动和非手动元素的时间和协调。初步研究表明,这些问题严重影响了手语动画的理解程度,目前的手语动画技术在这些方面需要改进。人类或动画角色的脸应该如何被清晰地表达出来,才能准确地表现出作为美国手语语法一部分的有语言意义的面部表情?这些动作的开始、偏移和过渡应该如何产生?面部表情和手部动作应该如何在时间上协调一致,以使美国手语的产生在语法上尽可能正确和可理解?为了回答诸如此类的开放性问题,pi的新方法将应用计算机视觉技术,对从人类手语者那里收集的带有语言注释的视频数据进行处理,以便生成用于动画制作的模型。pi将通过新的数据收集和注释来扩展他们现有的带注释的视频ASL语料库,并将分析这些数据来研究ASL生产中手动和非手动组件的使用、时间和同步。带注释的视频将用于训练高质量的计算机视觉模型,以识别语言上重要的面部表情和时间的微妙之处。这些计算机视觉模型的参数将用于假设美国手语时间和面部运动的计算模型,并将其纳入美国手语动画生成软件中,并由母语手语使用者进行评估。这些模型将在基于用户的研究周期中迭代改进,并纳入美国手语动画技术,以更准确地模仿人类手语。项目成果将包括用于美国手语表演动画的虚拟人物运动的高质量模型。对美国手语视频语料库的分析将对美国手语生产中的面部微表情和面部与手部的时间协调产生新的语言学见解,而对美国手语韵律分析的进展将有助于理解手语和口语之间的基本共性和特定情态差异,这对全面了解人类语言能力至关重要。新的建模方法和识别技术的创建将推动计算机视觉领域的发展,有利于识别和跟踪视频中快速和复杂的美国手语运动(以及其他形式的人类运动)中的人脸和身体。更广泛的影响:这项研究将导致产生语言准确的美国手语动画的技术的重大改进,这将使信息、应用程序、网站和服务更容易被大量英语读写能力相对较低的聋人获取。识别人类视频中手语的计算机视觉技术的进展将在人机交互、面部表情识别和动画以及计算机视觉方面具有普遍的适用性。该项目创建的语料库将使语言学和计算机科学的学生和研究人员(包括那些没有必要的技术和人力资源来收集母语签署者和时间密集的语言注释的人)能够从事美国手语的研究。要开发的技术还将使耗时的注释ASL视频语料库的创建部分自动化。与pi的早期工作一样,拟议中的研究将为聋哑人和其他未被充分代表的群体成员创造参与科学研究的机会。
英文摘要
American Sign Language (ASL) animations have the potential to make information accessible to many deaf adults in the United States who possess only limited English literacy. In this research, which involves collaboration across three institutions, the PIs' goal is to gain a better understanding of ASL linguistics through computational techniques while advancing the state of the art in the generation of ASL animations for accessibility applications for people who are deaf. To these ends, the PIs will develop linguistically based models of two aspects of ASL production: movements required for head gestures and facial expressions that carry essential grammatical information and frequently extend over domains larger than a single sign, and the timing and coordination of manual and non-manual elements of ASL signing. Preliminary work has shown that these issues significantly affect how well signers understand ASL animations, and that these aspects of current ASL animation technologies require improvement. How should the face of a human or animated character be articulated to perform, with accuracy, the linguistically meaningful facial expressions that are part of ASL grammar? How should the onsets, offsets, and transitions of these movements be produced? How should the facial expressions and hand movements be temporally coordinated so that the ASL production is as grammatically correct and understandable as possible? To answer open questions such as these, the PIs' novel approach will apply techniques from computer vision to linguistically annotated video data collected from human signers, in order to produce models for use in animation-production. The PIs will expand their existing annotated video ASL corpora through new data collection and annotation, and will analyze these data to study the use, timing, and synchronization of manual and non-manual components of ASL production. The annotated videos will be used to train high quality computer vision models for recognition of linguistically significant facial expressions and timing subtleties. Parameters of these computer vision models will be used to hypothesize computational models of ASL timing and facial movements, to be incorporated into ASL-animation generation software and evaluated by native signers. The models will be iteratively refined in cycles of user-based studies and incorporated into ASL animation technologies to more accurately mimic human signing. Project outcomes will include high quality models of the movement of virtual human characters for animations of ASL performance. The analysis of video corpora of ASL will produce new linguistic insights into the micro-facial expressions and the temporal coordination of the face and hands in ASL production, while advances in the analysis of ASL prosody will contribute to an understanding of the fundamental commonalities and modality-specific differences between signed and spoken languages that is essential to a full understanding of the human language faculty. The creation of new modeling approaches and recognition techniques will advance the field of computer vision, by benefiting the identification and tracking of the human face and body in video during the rapid and complex movements of ASL (and other forms of human movement).Broader Impacts: This research will lead to significant improvements to technology for generating linguistically accurate ASL animations, which will make information, applications, websites, and services more accessible to the large number of deaf individuals with relatively low English literacy. Advances in computer vision techniques for recognizing ASL in videos of humans will have general applicability in human-computer interaction, recognition and animation of facial expressions, and computer vision. The corpora created in this project will enable students and researchers in both linguistics and computer science (including those without access to the requisite technological and human resources to carry out their own data collection from native signers and time-intensive linguistic annotations) to engage in research on ASL. The techniques to be developed will also enable partial automation of the time-consuming creation of annotated ASL video corpora. As in the PIs' earlier work, the proposed research will create opportunities for people who are deaf and members of other underrepresented groups to participate in scientific research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Center: IUCRC Phase II Rutgers University: Center for Accelerated and Real Time Analytics (CARTA)
  • 批准号:
    2310966
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2023
  • 负责人:
    Dimitris Metaxas
  • 依托单位:
Collaborative Research: HCC: Medium: Linguistically-Driven Sign Recognition from Continuous Signing for American Sign Language (ASL)
  • 批准号:
    2212301
  • 项目类别:
    Standard Grant
  • 资助金额:
    $62.9万
  • 财政年份:
    2022
  • 负责人:
    Dimitris Metaxas
  • 依托单位:
NSF Convergence Accelerator Track H: AI-based Tools to Enhance Access and Opportunities for the Deaf
  • 批准号:
    2235405
  • 项目类别:
    Standard Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2022
  • 负责人:
    Dimitris Metaxas
  • 依托单位:
NSF Convergence Accelerator Track D: Data & AI Methods for Modeling Facial Expressions in Language with Applications to Privacy for the Deaf, ASL Education & Linguistic Res
  • 批准号:
    2040638
  • 项目类别:
    Standard Grant
  • 资助金额:
    $96.0万
  • 财政年份:
    2020
  • 负责人:
    Dimitris Metaxas
  • 依托单位:
海外基金