Teaching Future Big Data Analysts: Curriculum and Experience Report

Teaching Future Big Data Analysts: Curriculum and Experience Report
复制标题

教授未来大数据分析师:课程和经验报告

DOI:
--
复制
发表时间:
2017
期刊:
IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
J. Eckroth
J. Eckroth
中科院分区:
--
文献类型:
--
作者:
J. Eckroth

文献摘要

被引文献

相似文献

本文介绍了一所小型文理学院开设的“大数据挖掘与分析”课程的学习目标、课程设计、技术基础设施和课堂体验。该课程作为我们的数据分析未成年人的选修课,以及计算机科学和计算机信息系统专业的选修课。本课程向学生介绍数据分析,统计和使用Unix工具和R语言绘图。然后,它过渡到大数据项目,使用Apache Hadoop,HDFS和Map-Reduce; Apache Spark; Apache Hive;以及相关工具。一个主要的学习目标是让学生能够识别哪些工具最适合特定的数据集和数据分析任务。我们还希望学生能够将他们的发现传达给普通观众。作为潜在的未来数据分析师,我们的目标是为学生提供在未来职业生涯中有效解决数据分析问题(大数据或其他)的技能和敏感性。
This paper documents the learning objectives, curriculum design, technology infrastructure, and classroom experience for a "big data mining and analytics" course at a small liberal arts college. The course serves as an elective for our Data Analytics minor as well as an elective for computer science and computer information systems majors. The course introduces students to data analysis, statistics, and plotting with Unix tools and the R language. It then transitions into big data projects making use of Apache Hadoop, HDFS, and Map-Reduce; Apache Spark; Apache Hive; and related tools. A primary learning objective is that students demonstrate the ability to identify which tools are most appropriate for specific datasets and data analysis tasks. We also expect students to be able to communicate their findings to a general audience. As potential future data analysts, we aim to give students the skills and sensibility to efficiently solve data analysis problems, big data or otherwise, in their future careers.