Teaching Big Data with a Virtual Cluster

Teaching Big Data with a Virtual Cluster
复制标题

使用虚拟集群教授大数据

DOI:
10.1145/2839509.2844651
复制
发表时间:
2016
期刊:
Proceedings of the 47th ACM Technical Symposium on Computing Science Education
影响因子:
--
通讯作者:
J. Eckroth
J. Eckroth
中科院分区:
--
文献类型:
--
作者:
J. Eckroth

文献摘要

被引文献

相似文献

工业界和学术界都面临着大数据的挑战,即数据处理涉及的数据如此庞大或达到如此高的速度,以至于没有任何一台商用机器能够存储或处理所有数据。处理大数据的常见方法是将处理作业划分并分配给机器集群。理想情况下,教学生如何使用大数据的课程将为学生提供访问集群的机会以进行实践练习。然而,一组物理的本地机器可能非常昂贵,特别是在预算较小的小型机构中。在本报告中,我们总结了在一所小型私立文理学院的大数据挖掘和分析课程中开发和使用虚拟集群的经验。一台中等规模的服务器托管着一组虚拟机,这些虚拟机运行着流行的 Apache Hadoop 系统。虚拟集群为学生提供了实践经验,并且成本低于同等数量的物理机。它也很容易构建和重新配置。我们描述我们的实施,分析其性能特征,并将成本与物理集群和 Amazon Elastic MapReduce 云服务进行比较。我们总结了虚拟集群在课堂上的使用情况并展示了学生的反馈。
Both industry and academia are confronting the challenge of big data, i.e., data processing that involves data so voluminous or arriving at such high velocity that no single commodity machine is capable of storing or processing them all. A common approach to handling big data is to divide and distribute the processing job to a cluster of machines. Ideally, a course that teaches students how to work with big data would provide students access to a cluster for hands-on practice. However, a cluster of physical, on-premise machines may be prohibitively expensive, particularly at smaller institutions with smaller budgets. In this report, we summarize our experiences developing and using a virtual cluster in a big data mining and analytics course at a small private liberal arts college. A single moderately-sized server hosts a cluster of virtual machines, which run the popular Apache Hadoop system. The virtual cluster gives students hands-on experience and costs less than an equal number of physical machines. It is also easily constructed and reconfigured. We describe our implementation, analyze its performance characteristics, and compare costs with physical clusters and the Amazon Elastic MapReduce cloud service. We summarize our use of the virtual cluster in the classroom and show student feedback.