Resources for Computer-Based Sign Recognition from Video, and the Criticality of Consistency of Gloss Labeling across Multiple Large ASL Video Corpora

Resources for Computer-Based Sign Recognition from Video, and the Criticality of Consistency of Gloss Labeling across Multiple Large ASL Video Corpora
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
C. Neidle;Augustine Opoku;Carey M. Ballard;Konstantinos M. Dafnis;Evgenia Chroni;Dimitris N. Metaxas
C. Neidle;Augustine Opoku;Carey M. Ballard;Konstantinos M. Dafnis;Evgenia Chroni;Dimitris N. Metaxas
中科院分区:
其他
文献类型:
--
作者:
C. Neidle;Augustine Opoku;Carey M. Ballard;Konstantinos M. Dafnis;Evgenia Chroni;Dimitris N. Metaxas

文献摘要

被引文献

相似文献

WLASL声称是“单词级美国手语(ASL)识别的最大视频数据集”。它汇集了各种公开共享的视频集,这些视频集对标志识别研究非常有价值,并且已被广泛用于此类研究。然而,伴随的注释的一个关键问题迄今为止还没有被作者认识到,也没有被那些利用这些数据的人认识到:在符号产生和光泽标签之间没有1-1的对应关系。在这里,我们描述了一个大的(和最近扩大和增强),语言注释,可下载的,引用形式的ASL符号的视频语料库共享的美国手语语言研究项目(ASLLRP)-与23,452标志令牌和在线签署银行,在这种对应关系是强制执行的。此外,我们还为19,672个WLASL视频示例提供了符合ASLLRP注释约定的注释。对于那些希望使用WLASL视频的人来说,这提供了一组注释,使得有可能:(1)将这些数据可靠地用于计算研究;和/或(2)将WLASL和ASLLRP数据集联合收割机组合起来,创建一个比这些数据集中的任何一个单独更大和更丰富的组合资源,为所有标志提供一致的光泽标签。我们还提供了一个总结我们自己的标志识别研究的日期,利用这些数据资源。
The WLASL purports to be “the largest video dataset for Word-Level American Sign Language (ASL) recognition.” It brings together various publicly shared video collections that could be quite valuable for sign recognition research, and it has been used extensively for such research. However, a critical problem with the accompanying annotations has heretofore not been recognized by the authors, nor by those who have exploited these data: There is no 1-1 correspondence between sign productions and gloss labels. Here we describe a large (and recently expanded and enhanced), linguistically annotated, downloadable, video corpus of citation-form ASL signs shared by the American Sign Language Linguistic Research Project (ASLLRP)—with 23,452 sign tokens and an online Sign Bank—in which such correspondences are enforced. We furthermore provide annotations for 19,672 of the WLASL video examples consistent with ASLLRP glossing conventions. For those wishing to use WLASL videos, this provides a set of annotations that makes it possible: (1) to use those data reliably for computational research; and/or (2) to combine the WLASL and ASLLRP datasets, creating a combined resource that is larger and richer than either of those datasets individually, with consistent gloss labeling for all signs. We also offer a summary of our own sign recognition research to date that exploits these data resources.