Text Query based Traffic Video Event Retrieval with Global-Local Fusion Embedding

Text Query based Traffic Video Event Retrieval with Global-Local Fusion Embedding
复制标题

DOI:
10.1109/cvprw56347.2022.00353
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Thang-Long Nguyen-Ho;Minh Pham;Tien-Phat Nguyen;Hai-Dang Nguyen;M. Do;Tam V. Nguyen;M. Tran
Thang-Long Nguyen-Ho;Minh Pham;Tien-Phat Nguyen;Hai-Dang Nguyen;M. Do;Tam V. Nguyen;M. Tran
中科院分区:
其他
文献类型:
--
作者:
Thang-Long Nguyen-Ho;Minh Pham;Tien-Phat Nguyen;Hai-Dang Nguyen;M. Do;Tam V. Nguyen;M. Tran

文献摘要

相似文献

在快速发展的数据领域,基于文本描述的事件视频检索是一个很有前途的研究课题。然而,交通数据每天都在增加,因此需要智能交通系统管理与人类相结合,以加快搜索速度。我们提出了一个多模块系统,提供准确的结果,满足目标,包括可解释性和可扩展性在同一时间。我们的解决方案考虑与上述对象相关的邻居实体,以基于规则的方式表示事件,该方法可以通过多个对象的关系来表示事件。在我们提出的检索方法中,我们将修改后的阿里巴巴解决方案模型与AI City Challenge 2021中HCMUS方法的后处理技术相结合,以提高所获得结果的可解释性。由于交通数据以车辆为中心,我们采用语言和图像两个模块对输入数据进行分析,获得上下文的全局属性和车辆的内部属性。我们为每个表示向量引入了一对一的双重训练策略,以优化查询的内部特征。最后,精化模块收集先前的结果以增强最终的检索结果。我们以人工智能城市挑战赛2022的数据为基准,获得了MMR为0.3611的竞争结果。我们在50%的测试集中排名前4,在全部测试集中排名前5。
Retrieving event videos based on textual description is a promising research topic in the fast-growing data field. However, traffic data increases every day, so it is essential to need intelligent traffic system management in conjunction with humans to speed up the search. We propose a multi-module system that delivers accurate results that meet objectives, including explainability and scalability at the same time. Our solution considers neighbors entities related to the mentioned object to represent an event by rule-based, which can represent an event by the relationship of multiple objects. In our proposed retrieval method, we add our modified model of Alibaba solution with the post-processing techniques from HCMUS method in AI City Challenge 2021 to boost the explainability of the obtained results. As the traffic data is vehicle-centric, we apply two language and image modules to analyze the input data and obtain the global properties of the context and the internal attributes of the vehicle. We introduce a one-on-one dual training strategy for each representation vector to optimize the interior features for the query. Finally, a refinement module gathers previous results to enhance the final retrieval result. We benchmarked our approach on the data of the AI City Challenge 2022 and obtained the competitive results at an MMR of 0.3611. We were ranked in the top 4 on 50% of the test set and in the top 5 on the full set.