DiffSLVA: Harnessing Diffusion Models for Sign Language Video Anonymization

DiffSLVA: Harnessing Diffusion Models for Sign Language Video Anonymization
复制标题

DOI:
10.48550/arxiv.2311.16060
复制
发表时间:
2023-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhaoyang Xia;C. Neidle;Dimitris N. Metaxas
Zhaoyang Xia;C. Neidle;Dimitris N. Metaxas
中科院分区:
其他
文献类型:
--
作者:
Zhaoyang Xia;C. Neidle;Dimitris N. Metaxas

文献摘要

相似文献

由于美国手语(ASL)没有标准的书面形式,聋人签名者经常分享视频,以便用他们的母语进行交流。然而,由于手和脸都在手语中传达关键的语言信息,手语视频不能保护签名者的隐私。虽然签名者表示有兴趣在各种应用中使用手语视频匿名化,以有效保留语言内容,但鉴于手部动作和面部表情的复杂性,开发这种技术的尝试取得的成功有限。现有的方法主要依赖于视频片段中签名者的精确姿态估计,并且通常需要手语视频数据集进行训练。这些要求使他们无法在"野外“处理视频,部分原因是当前手语视频数据集的多样性有限。为了解决这些局限性,我们的研究引入了DiffSLVA,这是一种新的方法,它利用预训练的大规模扩散模型进行零拍摄文本引导的手语视频匿名化。我们采用ControlNet,它利用低级别的图像特征,如HED(整体嵌套边缘检测)边缘,以规避姿态估计的需要。此外,我们还开发了一个专门用于捕捉面部表情的模块,这对于用手语传达基本的语言信息至关重要。然后,我们结合联合收割机,实现匿名化,更好地保留了原始签名人的基本语言内容。这一创新方法首次使手语视频匿名化成为可能,可用于现实世界的应用,这将为聋人和听力障碍社区带来重大利益。我们证明了我们的方法的有效性与一系列的签名者匿名化实验。
Since American Sign Language (ASL) has no standard written form, Deaf signers frequently share videos in order to communicate in their native language. However, since both hands and face convey critical linguistic information in signed languages, sign language videos cannot preserve signer privacy. While signers have expressed interest, for a variety of applications, in sign language video anonymization that would effectively preserve linguistic content, attempts to develop such technology have had limited success, given the complexity of hand movements and facial expressions. Existing approaches rely predominantly on precise pose estimations of the signer in video footage and often require sign language video datasets for training. These requirements prevent them from processing videos 'in the wild,' in part because of the limited diversity present in current sign language video datasets. To address these limitations, our research introduces DiffSLVA, a novel methodology that utilizes pre-trained large-scale diffusion models for zero-shot text-guided sign language video anonymization. We incorporate ControlNet, which leverages low-level image features such as HED (Holistically-Nested Edge Detection) edges, to circumvent the need for pose estimation. Additionally, we develop a specialized module dedicated to capturing facial expressions, which are critical for conveying essential linguistic information in signed languages. We then combine the above methods to achieve anonymization that better preserves the essential linguistic content of the original signer. This innovative methodology makes possible, for the first time, sign language video anonymization that could be used for real-world applications, which would offer significant benefits to the Deaf and Hard-of-Hearing communities. We demonstrate the effectiveness of our approach with a series of signer anonymization experiments.