Prediction of Optimal Parallelism Level in Wide Area Data Transfers

Prediction of Optimal Parallelism Level in Wide Area Data Transfers
复制标题

广域数据传输中最佳并行度的预测

DOI:
--
复制
发表时间:
2011
影响因子:
5.3
通讯作者:
T. Kosar
T. Kosar
中科院分区:
计算机科学2区
文献类型:
--
作者:
E. Yildirim;Dengpan Yin;T. Kosar

文献摘要

被引文献

相似文献

广域数据传输可能是分布式应用程序端到端性能的主要瓶颈。在应用层增加广域吞吐量的一种实用方法是使用多个并行流。虽然增加并行流的数量可能比使用单个流产生更好的性能,但打开过多的流而使网络不堪重负可能会产生相反的效果。过多的流造成的拥塞可能导致吞吐量下降。因此,在不阻塞网络的情况下确定流的最佳数量是很重要的。预测这个“最优”数量并不简单,因为它取决于每个单独转移的许多特定参数。试图预测这一数字的通用模型要么过于依赖历史信息,要么无法实现准确的预测。在本文中,我们提出了一组新的模型,旨在以最少的历史信息和最低的预测开销来近似最优数量。提出了一种选择历史信息的最佳组合进行预测的算法,以达到评估的目的,并通过降低错误率来优化预测。通过使用很少的历史信息与实际的GridFTP数据传输进行比较,我们测量了所提出的预测模型的可行性和准确性,并且已经看到我们可以准确地预测并行流的吞吐量,并找到最优流数的非常接近的近似值。
Wide area data transfer may be a major bottleneck for the end-to-end performance of distributed applications. A practical way of increasing the wide area throughput at the application layer is using multiple parallel streams. Although increased number of parallel streams may yield much better performance than using a single stream, overwhelming the network by opening too many streams may have an inverse effect. The congestion created by excess number of streams may cause a drop down in the throughput achieved. Hence, it is important to decide on the optimal number of streams without congesting the network. Predicting this "optimum” number is not straightforward, since it depends on many parameters specific to each individual transfer. Generic models that try to predict this number either rely too much on historical information or fail to achieve accurate predictions. In this paper, we present a set of new models which aim to approximate the optimal number with least history information and lowest prediction overhead. An algorithm is introduced to select the best combination of historic information to do the prediction for evaluation purposes as well as optimizing prediction by reducing error rate. We measure the feasibility and accuracy of the proposed prediction models by comparing to actual GridFTP data transfer by using little historical information and have seen that we could predict the throughput of parallel streams accurately and find a very close approximation of the optimal stream number.