CyberTraining: CIU: SJSU Data Science for All Seminar Series
CyberTraining: CIU: SJSU Data Science for All Seminar Series
批准号:
1829622
负责人:
Leslie Albert
金额:
$41.01万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2022-08-31
中文摘要
国家的研究企业面临着数据科学家短缺的问题。扩大数据科学学生的渠道,特别是来自代表性不足的人群,需要教育机构提高对数据科学的认识,并在学生开始学术生涯时激发他们对数据的热情。目前,很少有社区学院或本科课程向广大学生提供网络基础设施工具或数据科学技术的培训。该项目采取了一种新的方法,通过培训社区学院和本科生,通过一系列“人人享有数据科学”的课外研讨会,为数据科学家提供数据分析支持,从而增强国家的数据科学劳动力。研讨会不需要事先掌握数据科学知识,强调可转移的技能,并为来自广泛学科和代表性不足的群体的学生提供了一条进入数据科学相关研究和其他职业的可行途径,而无需延长他们的毕业时间。通过提高国家的数据科学能力及其数据科学研究队伍的多样性,该项目符合国家利益,正如NSF的使命所述:促进科学进步,促进国家的繁荣和福利。该项目的目标是提高本科生对数据驱动科学的认识,并增加和多样化接受过数据处理培训的学生人数-准备数据进行分析所需的数据采集,转换,清理和分析。据业内专家介绍,数据争吵是数据科学的“重活”,占数据科学家日常工作的80%。将这种耗时的工作转移到训练有素的数据分析师身上,可以让数据科学家将更多的时间集中在研究上。该项目通过开发和提供广泛消费的课外研讨会来实现其目标,为本科生和社区大学生提供有关数据科学概念和行业领先的数据管理工具的互动培训。最初的研讨会主题是与项目顾问委员会合作选择的,包括Python、Apache Spark、Tableau和人工智能(AI)。研讨会对数据争论的关注也向学生介绍了数据准备文档-捕获可重复科学所需的数据出处。该项目对国家数据科学劳动力的贡献通过其研讨会材料和补充资源及其在线教师支持社区的免费和开放分发而扩大。为了鼓励在湾区社区学院和大学采用,通过共同指导和教师教学模式提供教师培训。该项目通过确定向不同人群的本科生教授数据科学的最有效教学方法,为教学研究做出贡献。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The Nation's research enterprise faces a shortage of data scientists. Expanding the pipeline of data science students, particularly from underrepresented populations, requires educational institutions to increase awareness of data science and inspire a passion for data in students as they begin their academic careers. Currently, few community colleges or undergraduate programs provide training in cyberinfrastructure tools or data science techniques to a broad student population. This project takes a novel approach to augmenting the Nation's data science workforce by training community college and undergraduate students to provide data analytics support to data scientists through a series of "Data Science for All" extracurricular seminars. The seminars require no prior data science knowledge, emphasize transferable skills, and present a feasible path into data science-related research and other careers for students from a broad array of disciplines and from underrepresented groups without extending their time to graduation. By increasing the Nation's data science capabilities and the diversity of its data science research workforce, the project serves the national interest, as stated by NSF's mission: to promote progress of science and advance the prosperity and welfare of the Nation. The goals of this project are to increase undergraduate student awareness of data-driven science and to grow and diversify the population of students trained to perform data wrangling - the data acquisition, transformation, cleaning, and profiling required to prepare data for analysis. According to industry experts, data wrangling is the "heavy lifting" of data science, constituting up to 80% of a data scientist's daily work. Shifting this time-consuming effort to trained data analysts free data scientists to focus more of their time on research. The project achieves its goals through the development and delivery of widely consumable, extracurricular seminars providing interactive training on data science concepts and industry-leading data wrangling tools to undergraduate and community college students. Initial seminar topics, selected in collaboration with the project's advisory board, include Python, Jupyter notebooks, Apache Spark, Tableau, and demystifying artificial intelligence (AI). The seminars' focus on data wrangling also introduces students to data preparation documentation - capturing the data provenance needed for reproducible science. This project's contribution to the Nation's data science workforce is broadened through the free and open distribution of its seminar materials and supplemental resources and its online instructor support community. To encourage adoption at Bay Area community colleges and universities, instructor training is provided through co-instruction and a teaching-the-teacher model. The project contributes to pedagogical research by identifying instructional approaches most effective in teaching data science to a diverse population of undergraduate students.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Data Science for All: Apache Spark & Jupyter Notebooks
全民数据科学:Apache Spark
DOI:
--
发表时间:
2021
期刊:
Pre-ICIS SIGDSA Symposium on Analytics and AI for a Sustainable and Resilient Future
影响因子:
--
作者:
[Jensen, Scott, Albert, Leslie, Huerta, Esperanza]
通讯作者:
Huerta, Esperanza
Python Foundations: Data Science for All
Python 基础:全民数据科学
DOI:
--
发表时间:
2019
期刊:
International Conference on Information Systems
影响因子:
--
作者:
[Albert, Leslie, Huerta, Esperanza, Jensen, Scott]
通讯作者:
Jensen, Scott
海外基金