WATSON: A Workflow-based Data Storage Optimizer for Analytics

WATSON: A Workflow-based Data Storage Optimizer for Analytics
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Jia Zou;Ming Zhao;Juwei Shi;Chen Wang
Jia Zou;Ming Zhao;Juwei Shi;Chen Wang
中科院分区:
其他
文献类型:
--
作者:
Jia Zou;Ming Zhao;Juwei Shi;Chen Wang

文献摘要

相似文献

- 本文研究一旦阅读了许多(蠕虫)方案,对数据间的数据放置参数进行了自动优化,其中首先通过生产者作业实现了数据,然后由一个或多个消费者工作进行了多次访问方案在大数据分析应用程序中无处不在,但现有的大数据自动调整技术通常集中在单个工作绩效上,以解决现有作品的缺点。研究数据放置参数有关阻止,分区和复制的模型,并通过生产者消费者模型对这些参数的不同配置引起的权衡,然后我们提出了一种新颖的跨层解决方案,该解决方案可以自动预测未来的工作负载的数据。访问模式和调整数据放置参数相应地优化了沃森的性能。各种分析工作负载加速。
—This paper studies the automatic optimization of data placement parameters for the inter-job write once read many (WORM) scenario where data is first materialized to storage by a producer job, and then accessed for many times by one or more consumer jobs. Such scenario is ubiquitous in Big Data analytics applications but existing Big Data auto-tuning techniques are often focused on single job performance. To address the shortcomings in existing works, this paper investigates data placement parameters regarding blocking, partitioning and replication and models the trade-offs caused by different configurations of these parameters through a producer-consumer model. We then present a novel cross-layer solution, WATSON, which can automatically predict future workloads’ data access patterns and tune data placement parameters accordingly to optimize the performance for an inter-job WORM scenario. WATSON can achieve up to eight times performance speedup on various analytics workloads.