Starfish: A Self-tuning System for Big Data Analytics

Starfish: A Self-tuning System for Big Data Analytics
复制标题

DOI:
--
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
H. Herodotou;Harold Lim;Gang Luo;N. Borisov;Liang Dong;Fatma Bilgen Cetin;S. Babu
H. Herodotou;Harold Lim;Gang Luo;N. Borisov;Liang Dong;Fatma Bilgen Cetin;S. Babu
中科院分区:
其他
文献类型:
--
作者:
H. Herodotou;Harold Lim;Gang Luo;N. Borisov;Liang Dong;Fatma Bilgen Cetin;S. Babu

文献摘要

被引文献

相似文献

对“大数据”的及时且具有成本效益的分析是在许多企业,科学和工程学科和政府努力中成功的关键人物。以及一系列声明性界面的过程是大多数大数据分析的大多数研究者。 Hadoop的表现有很大的需求,导致对资源,时间和金钱的使用(在Payas-you-go云中),我们介绍了一个用于大数据分析的自我调整系统Hadoop在适应用户需求和系统工作负载时自动提供良好的性能,而无需用户理解和操纵Hadoop中的许多调整旋钮。大数据的分析实践带来了新的挑战;
Timely and cost-effective analytics over “Big Data” is now a key ingredient for success in many businesses, scientific and engineering disciplines, and government endeavors. The Hadoop software stack—which consists of an extensible MapReduce execution engine, pluggable distributed storage engines, and a range of procedural to declarative interfaces—is a popular choice for big data analytics. Most practitioners of big data analytics—like computational scientists, systems researchers, and business analysts—lack the expertise to tune the system to get good performance. Unfortunately, Hadoop’s performance out of the box leaves much to be desired, leading to suboptimal use of resources, time, and money (in payas-you-go clouds). We introduce Starfish, a self-tuning system for big data analytics. Starfish builds on Hadoop while adapting to user needs and system workloads to provide good performance automatically, without any need for users to understand and manipulate the many tuning knobs in Hadoop. While Starfish’s system architecture is guided by work on self-tuning database systems, we discuss how new analysis practices over big data pose new challenges; leading us to different design choices in Starfish.