Isolated Sign Recognition using ASL Datasets with Consistent Text-based Gloss Labeling and Curriculum Learning

Isolated Sign Recognition using ASL Datasets with Consistent Text-based Gloss Labeling and Curriculum Learning
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Konstantinos M. Dafnis;Evgenia Chroni;C. Neidle;Dimitris N. Metaxas
Konstantinos M. Dafnis;Evgenia Chroni;C. Neidle;Dimitris N. Metaxas
中科院分区:
其他
文献类型:
--
作者:
Konstantinos M. Dafnis;Evgenia Chroni;C. Neidle;Dimitris N. Metaxas

文献摘要

被引文献

相似文献

我们提出了一种新的孤立手势识别方法,该方法结合了时空图形卷积网络(GCN)结构来建模人体骨架关键点,并对前后向视频流的后期融合进行了探索,并探索了课程学习的应用。我们采用了一种课程学习,它在训练期间动态地估计每个用于手势识别的输入视频的难度顺序;这涉及学习在训练期间动态更新的一组新的数据参数。这项研究利用了美国手语(ASL)的大型组合视频数据集,包括来自美国手语词典视频数据集(ASLLVD)和词级美国手语(WLASL)数据集的数据,并对后者进行了修改后的光泽标签-以确保光泽标签和不同的手语产品之间的1-1对应,以及两个数据集的光泽标签的一致性。这是首次将这两个数据集结合用于孤立手势识别研究。我们还比较了在组合数据集的几个不同子集上的手势识别性能,例如,每个手势的最小样本数(并且因此也在手势类和视频示例的总数上变化)。
We present a new approach for isolated sign recognition, which combines a spatial-temporal Graph Convolution Network (GCN) architecture for modeling human skeleton keypoints with late fusion of both the forward and backward video streams, and we explore the use of curriculum learning. We employ a type of curriculum learning that dynamically estimates, during training, the order of difficulty of each input video for sign recognition; this involves learning a new family of data parameters that are dynamically updated during training. The research makes use of a large combined video dataset for American Sign Language (ASL), including data from both the American Sign Language Lexicon Video Dataset (ASLLVD) and the Word-Level American Sign Language (WLASL) dataset, with modified gloss labeling of the latter—to ensure 1-1 correspondence between gloss labels and distinct sign productions, as well as consistency in gloss labeling across the two datasets. This is the first time that these two datasets have been used in combination for isolated sign recognition research. We also compare the sign recognition performance on several different subsets of the combined dataset, varying in, e.g., the minimum number of samples per sign (and therefore also in the total number of sign classes and video examples).