Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Exploratory analysis and performance prediction of big data transfer in High-performance Networks
复制标题

DOI:
10.1016/j.engappai.2021.104285
复制
发表时间:
2021-06
期刊:
Eng. Appl. Artif. Intell.
影响因子:
--
通讯作者:
Daqing Yun;Wuji Liu;C. Wu;N. Rao;R. Kettimuthu
Daqing Yun;Wuji Liu;C. Wu;N. Rao;R. Kettimuthu
中科院分区:
其他
文献类型:
--
作者:
Daqing Yun;Wuji Liu;C. Wu;N. Rao;R. Kettimuthu

文献摘要

被引文献

相似文献

大规模科学和商业应用中的大数据传输越来越多地通过高性能网络(HPN)中通过提前带宽预留提供的有保证带宽的连接进行。供应代理需要仔细安排数据传输请求、计算网络路径并分配适当的带宽。这种保留的带宽如果没有被充分利用,则可能由于在批准的时间窗口期间的独占访问而被简单地浪费,并且导致额外的开销和资源管理的复杂性。这就需要准确的性能预测,以保留符合实际需求的带宽,并避免过度配置。我们采用机器学习算法来预测大数据传输性能,这些算法基于过去几年收集的大量性能测量结果,这些测量结果是在几个真实的物理或模拟测试平台上使用不同的协议和工具包在不同的终端站点之间进行数据传输测试的结果。我们首先分析响应终端主机系统、网络连接和数据传输应用程序中的全面参数列表的性能模式,这些参数激发了机器学习的使用,并帮助我们识别潜在因素的影响。然后,我们提出了基于阈值和聚类的方法来消除数据预处理中潜在因素的负面影响,并基于定制的面向领域的损失函数构建一个强大的性能预测器。所提出的方法的性能进行了验证,通过广泛的实验,使用SVR和RFR以及一般性能界的理论分析。
Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.