Including Facial Expressions in Contextual Embeddings for Sign Language Generation

Including Facial Expressions in Contextual Embeddings for Sign Language Generation
复制标题

DOI:
10.18653/v1/2023.starsem-1.1
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Carla Viegas;Mert Inan;Lorna C. Quandt;Malihe Alikhani
Carla Viegas;Mert Inan;Lorna C. Quandt;Malihe Alikhani
中科院分区:
其他
文献类型:
--
作者:
Carla Viegas;Mert Inan;Lorna C. Quandt;Malihe Alikhani

文献摘要

相似文献

现有的手语生成框架缺乏表情性和自然性,这是由于只关注手势,忽视了面部表情的情感、语法和语义功能。这项工作的目的是通过基础面部表情来增强手语的语义表示。我们研究的效果建模的文字,光泽和面部表情之间的关系的性能的标志生成系统。特别是,我们提出了一个双编码器Transformer能够生成手动标志,以及面部表情,通过捕获的相似性和差异,发现在文本和标志光泽注释。我们考虑到面部肌肉活动在表达手语强度方面的作用,率先在手语生成中采用面部动作单元。我们进行了一系列的实验表明,我们提出的模型提高了自动生成的手语的质量。
State-of-the-art sign language generation frameworks lack expressivity and naturalness which is the result of only focusing manual signs, neglecting the affective, grammatical and semantic functions of facial expressions. The purpose of this work is to augment semantic representation of sign language through grounding facial expressions. We study the effect of modeling the relationship between text, gloss, and facial expressions on the performance of the sign generation systems. In particular, we propose a Dual Encoder Transformer able to generate manual signs as well as facial expressions by capturing the similarities and differences found in text and sign gloss annotation. We take into consideration the role of facial muscle activity to express intensities of manual signs by being the first to employ facial action units in sign language generation. We perform a series of experiments showing that our proposed model improves the quality of automatically generated sign language.