A performance study of big data analytics platforms
A performance study of big data analytics platforms
复制标题
DOI:
10.1109/bigdata.2017.8258260
复制
发表时间:
2017-12
期刊:
影响因子:
--
通讯作者:
Pouria Pirzadeh;M. Carey;T. Westmann
中科院分区:
文献类型:
--
作者:
Pouria Pirzadeh;M. Carey;T. Westmann
Big Data analytics has become an invaluable tool in a wide variety of businesses for exploiting the wealth of Big Data that they now have access to. As a result, various solutions within different categories of Big Data systems are emerging to meet their needs. In this paper we use the TPC-H benchmark to compare the performance of four Big Data systems picked from the major categories of Big Data platforms: a commercial parallel relational database (from the traditional DBMS world), Hive and Spark SQL (from the SQL-on-Hadoop world), and AsterixDB (from the world of NoSQL systems). All of these systems have sufficiently rich query APIs and runtime systems to run TPC-H in its full form. On the other hand, the systems also have major differences in terms of their architectures, preferred storage formats, support for complex schema definitions, and approaches to query processing. This makes them a very interesting set of representative Big Data systems to compare. We present the results that we obtained through running these systems at different TPC-H scales using various settings, and we analyze a selected set of interesting query results in more detail to explore the trade-offs between performance, storage formats, and schema definitions. A follow-up discussion is included as well to summarize the lessons learned from this effort.