OneFi: One-Shot Recognition for Unseen Gesture via COTS WiFi

OneFi: One-Shot Recognition for Unseen Gesture via COTS WiFi
复制标题

DOI:
10.1145/3485730.3485936
复制
发表时间:
2021-11
期刊:
Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems
影响因子:
--
通讯作者:
Rui Xiao;Jianwei Liu;Jinsong Han;K. Ren
Rui Xiao;Jianwei Liu;Jinsong Han;K. Ren
中科院分区:
其他
文献类型:
--
作者:
Rui Xiao;Jianwei Liu;Jinsong Han;K. Ren

文献摘要

被引文献

相似文献

基于 WiFi 的人体手势识别 (HGR) 在无设备人机交互方面变得越来越有前景。然而,由于可扩展性有限,尤其是对于看不见的手势,现有的基于 WiFi 的方法尚未准备好用于实际部署。其背后的原因是,在引入看不见的手势时,先前的工作必须收集大量样本并重新训练模型。虽然最近few-shot学习的进展为解决这个问题带来了新的机会,但开销并没有得到有效降低。这是因为这些方法仍然需要大量数据来学习足够的先验知识,并且其复杂的训练过程增加了常规训练成本。在本文中,我们提出了一种基于 WiFi 的 HGR 系统,即 OneFi,它只需一个(或几个)标记样本即可识别看不见的手势。 OneFi 从根本上解决了高开销的挑战。一方面,OneFi利用虚拟手势生成机制,使得数据收集过程中的大量工作可以得到显着减轻。另一方面,OneFi 采用基于转导微调的轻量级一次性学习框架来消除模型重新训练。我们还设计了一个基于自注意力的主干网,称为 WiFi Transformer,以最大限度地减少所提出框架的训练成本。我们使用商用 WiFi 设备建立了一个真实世界的测试平台,并对其进行了广泛的实验。评估结果表明,在有1、3、5、7个标记样本的情况下,OneFi能够识别看不见的手势,准确率分别为84.2%、94.2%、95.8%和98.8%,而整个训练过程不到两分钟。
WiFi-based Human Gesture Recognition (HGR) becomes increasingly promising for device-free human-computer interaction. However, existing WiFi-based approaches have not been ready for real-world deployment due to the limited scalability, especially for unseen gestures. The reason behind is that when introducing unseen gestures, prior works have to collect a large number of samples and re-train the model. While the recent advance of few-shot learning has brought new opportunities to solve this problem, the overhead has not been effectively reduced. This is because these methods still require enormous data to learn adequate prior knowledge, and their complicated training process intensifies the regular training cost. In this paper, we propose a WiFi-based HGR system, namely OneFi, which can recognize unseen gestures with only one (or few) labeled samples. OneFi fundamentally addresses the challenge of high overhead. On the one hand, OneFi utilizes a virtual gesture generation mechanism such that the massive efforts in prior works can be significantly alleviated in the data collection process. On the other hand, OneFi employs a lightweight one-shot learning framework based on transductive fine-tuning to eliminate model re-training. We additionally design a self-attention based backbone, termed as WiFi Transformer, to minimize the training cost of the proposed framework. We establish a real-world testbed using commodity WiFi devices and perform extensive experiments over it. The evaluation results show that OneFi can recognize unseen gestures with the accuracy of 84.2, 94.2, 95.8, and 98.8% when 1, 3, 5, 7 labeled samples are available, respectively, while the overall training process takes less than two minutes.