Unsupervised Key Hand Shape Discovery of Sign Language Videos with Correspondence Sparse Autoencoders

Unsupervised Key Hand Shape Discovery of Sign Language Videos with Correspondence Sparse Autoencoders
复制标题

使用对应稀疏自动编码器进行手语视频的无监督关键手形发现

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
L. Akarun
L. Akarun
中科院分区:
--
文献类型:
--
作者:
Recep Doga Siyli;Batuhan Gündogdu;M. Saraçlar;L. Akarun

文献摘要

被引文献

相似文献

手语识别是一项困难的任务,通常需要手语专家进行繁琐的注释。绕过帧级注释的端到端学习尝试在有限的数据集中取得了一些成功,但已经证明高质量的注释可以大幅提高性能。最近使用深度神经网络的无监督学习方法在学习特征提取方面取得了成功。然而,不存在使用无监督技术的高质量帧级分类的技术。在本文中,我们使用端到端神经网络架构分配一个孤立的手语(SL)数据集的标签,这些架构在语音处理中的子词声学单元的无监督发现方面已经证明是成功的。我们观察到,关键手形状(KHS),这是有意义的视觉基本部分的迹象在SL数据集可以检测使用无监督聚类技术。稀疏自动编码器可以成功地检索和集群KHS中使用的孤立的迹象。此外,在自动编码器方案中使用对应帧具有继续学习过程的能力。
Recognition of sign language is a difficult task which often requires tedious annotations by sign language experts. End-to-end learning attempts that bypass frame level annotations have achieved some success in limited datasets, but it has been shown that high quality annotations improve performance drastically. Recent unsupervised learning methods using deep neural networks have achieved successes in learning feature extraction. Yet a technique for high quality frame level classification using unsupervised techniques does not exist. In this paper, we assign labels of an isolated Sign Language(SL) dataset using end-to-end neural network architectures that have proven success in unsupervised discovery of sub-word acoustic units in speech processing. We observe that key-hand-shape s(KHS), which are meaningful visual basic parts of signs in a SL dataset can be detected using unsupervised clustering techniques. Sparse autoencoders can successfully retrieve and cluster KHS s used in isolated signs. In addition, using correspondent frames in an autoencoder scheme has the power to continue the learning process.