ViTag: Online WiFi Fine Time Measurements Aided Vision-Motion Identity Association in Multi-person Environments

ViTag: Online WiFi Fine Time Measurements Aided Vision-Motion Identity Association in Multi-person Environments
复制标题

DOI:
10.1109/secon55815.2022.9918171
复制
发表时间:
2022-09
期刊:
2022 19th Annual IEEE International Conference on Sensing, Communication, and Networking (SECON)
影响因子:
--
通讯作者:
Bryan Bo Cao;Abrar Alali;Hansi Liu;Nicholas Meegan;M. Gruteser;Kristin J. Dana;A. Ashok;Shubham Jain
Bryan Bo Cao;Abrar Alali;Hansi Liu;Nicholas Meegan;M. Gruteser;Kristin J. Dana;A. Ashok;Shubham Jain
中科院分区:
其他
文献类型:
--
作者:
Bryan Bo Cao;Abrar Alali;Hansi Liu;Nicholas Meegan;M. Gruteser;Kristin J. Dana;A. Ashok;Shubham Jain

文献摘要

相似文献

在本文中,我们提出了ViTag来关联用户身份跨多模态数据,特别是那些从相机和智能手机获得的数据。ViTag将一系列视觉跟踪器生成的边界框与来自智能手机的惯性测量单元(IMU)数据和Wi-Fi精细时间测量(FTM)相关联。我们将问题表述为序列到序列(seq2seq)转换的关联。在这个两步过程中,我们的系统首先使用多模态LSTM编码器-解码器网络(X-Translator)执行跨模态转换,该网络将一种模态转换为另一种模态,例如,纯粹从相机边界盒重建IMU和FTM读数。其次,关联模块查找相机和手机域之间的身份匹配,然后将翻译的模态与来自同一模态的观测数据进行匹配。与现有的工作相反,我们提出的方法可以在所有用户可能执行相同活动的多人场景中关联身份。在真实的室内和室外环境中进行的大量实验表明,相机和手机数据(IMU和FTM)的在线关联在1到3秒的窗口内实现了88.39%的平均身份精度(IDP),优于最先进的Vi-Fi(82.93%)。对电话域内模态的进一步研究表明,FTM可以平均提高12.56%的关联性能。最后,通过灵敏度实验验证了ViTag在不同噪声和环境变化下的鲁棒性。
In this paper, we present ViTag to associate user identities across multimodal data, particularly those obtained from cameras and smartphones. ViTag associates a sequence of vision tracker generated bounding boxes with Inertial Mea-surement Unit (IMU) data and Wi-Fi Fine Time Measurements (FTM) from smartphones. We formulate the problem as association by sequence to sequence (seq2seq) translation. In this two-step process, our system first performs cross-modal translation using a multimodal LSTM encoder-decoder network (X-Translator) that translates one modality to another, e.g. recon-structing IMU and FTM readings purely from camera bounding boxes. Second, an association module finds identity matches between camera and phone domains, where the translated modality is then matched with the observed data from the same modality. In contrast to existing works, our proposed approach can associate identities in multi-person scenarios where all users may be performing the same activity. Extensive experiments in real-world indoor and outdoor environments demonstrate that online association on camera and phone data (IMU and FTM) achieves an average Identity Precision Accuracy (IDP) of 88.39% on a 1 to 3 seconds window, outperforming the state-of-the-art Vi-Fi (82.93%). Further study on modalities within the phone domain shows the FTM can improve association performance by 12.56% on average. Finally, results from our sensitivity experiments demonstrate the robustness of ViTag under different noise and environment variations.