File-access patterns of data-intensive workflow applications and their implications to distributed filesystems

File-access patterns of data-intensive workflow applications and their implications to distributed filesystems
复制标题

DOI:
10.1145/1851476.1851585
复制
发表时间:
2010-06
期刊:
--
影响因子:
--
通讯作者:
Takeshi Shibata;SungJun Choi;K. Taura
Takeshi Shibata;SungJun Choi;K. Taura
中科院分区:
其他
文献类型:
--
作者:
Takeshi Shibata;SungJun Choi;K. Taura

文献摘要

被引文献

相似文献

本文研究了自然语言处理、天文图像分析和web数据分析等五个实际数据密集型工作流应用。数据密集型工作流正日益成为集群和网格环境中的重要应用。它们对工作流执行环境的各种组件(包括作业调度器、调度器、文件系统和文件分级工具)提出了新的挑战。实现高性能的关键是在执行主机之间进行有效的数据共享,以及减少数据传输量的位置感知调度。虽然在调度工作流方面已经做了很多工作,但其中许多工作使用合成的或随机的工作负载。因此,它们对实际工作负载的影响在很大程度上是未知的。了解现实世界工作流应用程序的特征是促进这一领域研究的必要步骤。为此,我们分析了现实世界中的工作流应用程序,重点关注它们的文件访问模式,并总结了它们对调度器和文件系统/分级设计的影响。
This paper studies five real-world data intensive workflow applications in the fields of natural language processing, astronomy image analysis, and web data analysis. Data intensive workflows are increasingly becoming important applications for cluster and Grid environments. They open new challenges to various components of workflow execution environments including job dispatchers, schedulers, file systems, and file staging tools. The keys to achieving high performance are efficient data sharing among executing hosts and locality-aware scheduling that reduces the amount of data transfer. While much work has been done on scheduling workflows, many of them use synthetic or random workload. As such, their impacts on real workloads are largely unknown. Understanding characteristics of real-world workflow applications is a required step to promote research in this area. To this end, we analyse real-world workflow applications focusing on their file access patterns and summarize their implications to schedulers and file system/staging designs.