BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues

BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues
复制标题

DOI:
10.1007/978-3-030-58621-8_3
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Samuel Albanie;Gül Varol;Liliane Momeni;Triantafyllos Afouras;Joon Son Chung;Neil Fox;Andrew Zisserman
Samuel Albanie;Gül Varol;Liliane Momeni;Triantafyllos Afouras;Joon Son Chung;Neil Fox;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Samuel Albanie;Gül Varol;Liliane Momeni;Triantafyllos Afouras;Joon Son Chung;Neil Fox;Andrew Zisserman

文献摘要

相似文献

最近在细粒度手势和动作分类以及机器翻译方面的进展表明,自动手语识别成为现实的可能性。实现这一目标的关键障碍是缺乏适当的训练数据,这源于标识标注的高度复杂性和合格标注人员的有限供应。在这项工作中,我们引入了一种新的可扩展的方法来收集连续视频中的符号识别数据。我们利用弱对齐字幕的广播素材和关键字识别方法来自动定位1,000小时视频中1,000个标志的词汇表。我们做出了以下贡献:(1)我们展示了如何使用手语的口型线索从视频数据中获得高质量的注释-结果是BSL- 1k数据集,这是一个空前规模的英国手语(BSL)符号的集合;(2)我们表明,我们可以使用BSL- 1k来训练BSL中共同表达符号的强符号识别模型,并且这些模型还为其他手语和基准形成了出色的预训练-我们在MSASL和WLASL基准上都超过了最先进的水平。最后,(3)我们提出了用于标识识别和标识识别任务的新的大规模评估集,并提供了基线,我们希望这将有助于促进该领域的研究。
Recent progress in fine-grained gesture and action classification, and machine translation, point to the possibility of automated sign language recognition becoming a reality. A key stumbling block in making progress towards this goal is a lack of appropriate training data, stemming from the high complexity of sign annotation and a limited supply of qualified annotators. In this work, we introduce a new scalable approach to data collection for sign recognition in continuous videos. We make use of weakly-aligned subtitles for broadcast footage together with a keyword spotting method to automatically localise sign-instances for a vocabulary of 1,000 signs in 1,000 h of video. We make the following contributions: (1) We show how to use mouthing cues from signers to obtain high-quality annotations from video data—the result is the BSL-1K dataset, a collection of British Sign Language (BSL) signs of unprecedented scale; (2) We show that we can use BSL-1K to train strong sign recognition models for co-articulated signs in BSL and that these models additionally form excellent pretraining for other sign languages and benchmarks—we exceed the state of the art on both the MSASL and WLASL benchmarks. Finally, (3) we propose new large-scale evaluation sets for the tasks ofsign recognitionandsign spottingand provide baselines which we hope will serve to stimulate research in this area.