How GridFTP Pipelining, Parallelism and Concurrency Work: A Guide for Optimizing Large Dataset Transfers

How GridFTP Pipelining, Parallelism and Concurrency Work: A Guide for Optimizing Large Dataset Transfers
复制标题

GridFTP 管道、并行性和并发性如何工作:优化大型数据集传输指南

DOI:
10.1109/sc.companion.2012.73
复制
发表时间:
2012
期刊:
2012 SC Companion: High Performance Computing, Networking Storage and Analysis
影响因子:
--
通讯作者:
T. Kosar
T. Kosar
中科院分区:
--
文献类型:
--
作者:
E. Yildirim;Jangyoung Kim;T. Kosar

文献摘要

被引文献

相似文献

优化大文件在高带宽网络上的传输是一项具有挑战性的任务,需要考虑许多参数(例如网络速度、往返时间和当前流量)。不幸的是,当传输由许多小文件组成的数据集时,此任务变得更加复杂。在这种情况下,大数据集传输的性能不仅取决于传输协议和网络的特性,还取决于构成数据集的文件的数量和大小分布。GridFTP是最先进的传输工具,它提供了克服大型数据集传输瓶颈的功能。GridFTP的三个最重要的参数是流水线、并行性和并发性。在这项研究中,我们研究了这三个重要参数的影响,这些参数的优化提供了模型,定义的指导方针,并给出了一个算法,为他们的实际使用不同大小的文件的大型数据集的传输。
Optimizing the transfer of large files over high-bandwidth networks is a challenging task that requires the consideration of many parameters (e.g. network speed, roundtrip time, and current traffic). Unfortunately, this task becomes more complex when transferring datasets comprised of many small files. In this case, the performance of large dataset transfers not only depends on the characteristics of the transfer protocol and network, but also the number and the size distribution of the files that constitute the dataset. GridFTP is the most advanced transfer tool that provides functions to overcome large dataset transfer bottlenecks. Three of the most important parameters of GridFTP are pipelining, parallelism and concurrency. In this study, we research the effects of these three important parameters, provide models for optimization of these parameters, define guidelines and give an algorithm for their practical use for transfer of large datasets of varying size files.