2016 Ieee International Conference on Big Data (big Data) towards Optimizing Large-scale Data Transfers with End-to-end Integrity Verification

2016 Ieee International Conference on Big Data (big Data) towards Optimizing Large-scale Data Transfers with End-to-end Integrity Verification
复制标题

2016 年 IEEE 大数据国际会议通过端到端完整性验证优化大规模数据传输

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
M. Papka
M. Papka
中科院分区:
--
文献类型:
--
作者:
Si Liu;Eun;R. Kettimuthu;Xian;M. Papka

文献摘要

被引文献

相似文献

——由实验设备和高性能计算设备模拟产生的科学数据规模迅速增长。在许多情况下,这些数据需要快速可靠地传输到远程设施进行存储、分析、共享等。同时,用户希望在数据被写入目标磁盘后通过校验和来验证数据的完整性,以确保文件没有被损坏,例如由于网络或存储数据损坏、软件错误或人为错误。这种端到端完整性验证会产生额外的开销(额外的磁盘I/O和更多的计算),并增加总体数据传输时间。在本文中,我们评估了最大化数据传输和校验和计算之间重叠的策略。更具体地说,我们评估文件级和块级(具有不同块大小)管道来重叠数据传输和校验和计算。我们在GridFTP(一种广泛用于科学数据传输的协议)的背景下评估这些管道方法。我们进行了理论分析和实际实验来评估我们的方法。结果表明,块级流水线是最大化数据传输和校验和计算之间重叠的有效方法,与连续执行传输和校验和相比,可以将端到端完整性验证的总体数据传输时间提高70%,与文件级流水线相比可提高60%。
—The scale of scientific data generated by experimental facilities and simulations on high-performance computing facilities has been growing rapidly. In many cases, this data needs to be transferred rapidly and reliably to remote facilities for storage, analysis, sharing etc. At the same time, users want to verify the integrity of the data by doing a checksum after the data has been written to disk at the destination, to ensure the file has not been corrupted, for example due to network or storage data corruption, software bugs or human error. This end-to-end integrity verification creates additional overhead (extra disk I/O and more computation) and increases the overall data transfer time. In this paper, we evaluate strategies to maximize the overlap between data transfer and checksum computation. More specifically, we evaluate file-level and block-level (with various block sizes) pipelining to overlap data transfer and checksum computation. We evaluate these pipelining approaches in the context of GridFTP, a widely used protocol for science data transfers. We conducted both theoretical analysis and real experiments to evaluate our methods. The results show that block-level pipelining is an effective method in maximizing the overlap between data transfer and checksum computation and can improve the overall data transfer time with end-to-end integrity verification by up to 70% compared to the sequential execution of transfer and checksum, and by up to 60% compared to file-level pipelining.