Optimization of Big Data Parallel Scheduling Based on Dynamic Clustering Scheduling Algorithm

Optimization of Big Data Parallel Scheduling Based on Dynamic Clustering Scheduling Algorithm
复制标题

DOI:
10.1007/s11265-022-01765-4
复制
发表时间:
2022-06
期刊:
Journal of Signal Processing Systems
影响因子:
--
通讯作者:
F. Liu;Yanxiang He;Jing He;Xing Gao;Feihu Huang
F. Liu;Yanxiang He;Jing He;Xing Gao;Feihu Huang
中科院分区:
其他
文献类型:
--
作者:
F. Liu;Yanxiang He;Jing He;Xing Gao;Feihu Huang

文献摘要

相似文献

在当今的数据时代,随着海量数据的增加,大数据处理分析框架在海量信息处理中发挥着重要的作用。数据共享是为了通过结构化数据调度来提升数据处理的性能。然而,这种方法使得额外的数据复制和缓存的通信成本和缓存成本更高。因此,在大数据分析环境下,本文利用基于数据相关性的动态集群调度算法(DCSA)对大数据任务进行并行优化。首先,生成基于服务器请求数据库的动态数据队列。将数据项的优先级和数据项的大小作为动态数据队列进行数据聚类关联的考虑因素。然后引入权值,对动态数据项进行均衡化,为多通道优化调度提供依据。其次,根据数据项的相关性,采用数据优化放置机制,将聚集在同一帧中的数据进行优化放置。在放置完成后,以数据项的局部特性为约束,对动态数据进行统一调度,使迁移时的代价最小。通过目标迭代调整最优调度方案,最终实现多通道最优调度。实验表明,该方法能够实现动态数据的最优调度。
In today’s data age, the big data processing analysis framework plays an important role in mass information processing, along with the increasing of massive data. “Sharing Data” is proposed to enhance the performance of data processing through structured data scheduling. However, such approach makes the higher communication cost and buffer cost for the extra data copy and buffering. Hence, in the big data analysis environment, this paper uses based on the correlation of data, Dynamic Cluster Scheduling Algorithm(DCSA) is proposed for parallel optimization of big data tasks. Firstly, a dynamic data queue based on the server’s request database is generated. The priority of data item and size of data item are as the considerations of dynamic data queue for data clustering association. And then the weights are introduced, the dynamic data item is made equalization to provide the basis for the multi-channel optimal scheduling. Secondly, according to the relevance of the data items, the mechanism of data optimized placement is used to make the data which are aggregated in the same frame. After the placement is completed, the dynamic data is uniformly scheduled to minimize the cost at the time of migration, with the local characteristics of the data item as constraints. Through the target iteration, the optimal scheduling scheme is adjusted, and finally to achieve multi-channel optimal scheduling. Experiments show that the proposed method enables dynamic data to achieve optimal scheduling.