A cost-efficient framework for crowdsourced data collection in vehicular networks

A cost-efficient framework for crowdsourced data collection in vehicular networks
复制标题

车辆网络中众包数据收集的经济高效框架

DOI:
10.1109/jiot.2021.3065716
复制
发表时间:
2021
影响因子:
10.6
通讯作者:
Jiazhuang Lu
Jiazhuang Lu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bo Yin;Jiazhuang Lu

文献摘要

相似文献

车载网络因其强大的感知能力和强大的移动性而被公认为是一种创新的信息收集技术,已成为地理众包服务的重要平台,在环境数据收集方面尤其有用。车辆网络中的众包数据收集由于资源受限的无线通信链路而必须以通信高效的方式执行。此外,由于工作人员可能提供质量差的数据,甚至伪造数据以骗取奖励,因此提高响应质量是一个紧迫的问题。虽然现有的工作集中在空间/时间任务覆盖和金钱奖励,在这项工作中,我们提出了一个具有成本效益的框架,在车辆网络中的众包数据收集,同时确保响应的准确性。众包数据收集包括两个主要步骤:1)任务分配和2)众包答案收集。我们首先提出了一个任务分配方案,最大限度地提高整体数据质量,减少数据传输量。我们基于高斯混合模型估计数据质量水平,并通过仔细选择用于众包任务的车辆子集来减少数据传输量。然后,我们设计了一个答案收集计划,同时考虑聚合树的长度和数据传输,并最大限度地减少从参与者收集答案的通信成本。在合成数据集和真实的数据集上的实验表明,我们提出的框架取得了令人满意的结果。
Vehicular networks, which are recognized as an innovative technology for information collection due to the powerful sensing capability and strong mobility, have been an important platform for geographic crowdsourcing services and are particularly useful in environmental data collection. Crowdsourced data collection in vehicular networks must be performed in a communication-efficient manner due to the resource-constrained wireless communication links. Moreover, it is a pressing problem to improve the quality of responses because workers may provide data of poor quality or even fabricate data to defraud rewards. While the existing work has focused on spatial/temporal task coverage and monetary rewards, in this work, we propose a cost-efficient framework for crowdsourced data collection in vehicular networks while ensuring the accuracy of the response. Crowdsourced data collection consists of two main steps: 1) task assignment and 2) crowdsourced answer gathering. We first propose a task assignment scheme that maximizes the overall data quality and reduces the amount of data transmission. We estimate the data quality level based on a Gaussian mixture model and reduce the amount of data transmission by carefully selecting a subset of vehicles for crowdsourced tasks. We then design an answer gathering scheme that considers both the length of the aggregation tree and data delivery and minimizes the communication cost for collecting answers from participants. Extensive experiments on both synthetic data sets and real data sets show that our proposed framework achieves promising results.