课题基金 / 基金详情

项目摘要

项目成果

Jeffrey T. Leek的其他基金

相似基金

相关文献

中文摘要
翻译
 描述(由申请人提供):Sciencefic结果存在可重复性和可复制性的危机。无论是在《科学》杂志(Sciencefic)还是在《白杨》杂志上,这场危机日益引起人们的关注。这场危机如此严重,以至于美国国会目前正在调查Sciencefic过程的可重复性。这场危机的核心是整个Sciencefic企业缺乏数据分析技能。一种正在形成的共识是,解决这场危机的最佳方式是增加数据分析培训,特别是关于可再现性和可复制性的培训。在这项应用中,我们(1)提出了可重复性和可复制性的fi第一正式统计模型,然后使用世界上最大的大规模在线开放计划的数据和实验来(2)进行随机研究,以改善我们对哪些统计方法和协议导致普通用户手中的可重复性和可复制性增加的知识,以及(3)分析学习者、课程和内容特征,以提高学习者的成功率和吞吐量,以增加全球训练有素的数据分析师的数量。为了实现目标(2)和(3),我们将使用世界上规模最大、吞吐量最高的数据科学计划:约翰·霍普金斯数据科学专业化认证。该专业由该项目的研究人员开发,包括每月提供的九门课程。自2014年4月该项目启动以来,这些课程的注册人数已超过200万人,他们几乎所有的经历都被记录为数据。此外,本系列的MOOC平台允许随机分配测验问题和内容。我们将通过开源软件、分析协议、我们广受欢迎的博客和数据科学专业化认证来传播我们的成果,以最大限度地改进数据科学培训,并减少科学fic复制和再现性问题。该项目的规模意味着,通过提高项目的质量和完成课程的学生数量,即使是很小的百分比,我们也可以影响全球数据分析行为。
英文摘要
 DESCRIPTION (provided by applicant): There is a crisis of reproducibility and replicability of scientific results. This crisis is an increasing source of concern both in the scientific and poplar press. The crisis is so acute that the United States Congress is currently investigating reproducibility of the scientific process. At the heart of the crisis is a shortage of data analytc skill throughout the scientific enterprise. There is an emerging consensus that the best way to address the crisis is to increase data analytic training, particularly around reproducibility and replicability. In this application we (1) propose the first formal statistical model for reproduciility and replicability and then use data and experiments from the largest massive online open program in data science in the world to (2) perform randomized studies to improve our knowledge about which statistical methods and protocols lead to increased reproducibility and replicability in the hands of average users and (3) to analyze learner, course, and content characteristics that increase learner success and throughput to increase the number of trained data analysts worldwide. To accomplish goals (2) and (3) we will use the largest and highest throughput data science program in the world: the Johns Hopkins Data Science Specialization. This specialization, developed by the investigators of this project, consists of nine courses that are offered every month. Since the launch of this program in April 2014, these classes have seen more than two million enrollments and nearly all their experiences have been recorded as data. Furthermore, the MOOC platform for this series permits random assignment of quiz questions and content. We will disseminate our results through open source software, analysis protocols, our popular blog, and the Data Science Specialization to maximally improve data science training and reduce the scientific replication and reproducibility problem. The size of ths program means that by increasing quality of the program and the number of completing students by even a small percentage we can affect global data analytic behavior.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
  • 批准号:
    10598130
  • 项目类别:
  • 资助金额:
    $42.82万
  • 财政年份:
    2022
  • 负责人:
    Jeffrey T. Leek
  • 依托单位:
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
  • 批准号:
    10330636
  • 项目类别:
  • 资助金额:
    $2.68万
  • 财政年份:
    2022
  • 负责人:
    Jeffrey T. Leek
  • 依托单位:
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
  • 批准号:
    10654376
  • 项目类别:
  • 资助金额:
    $40.44万
  • 财政年份:
    2022
  • 负责人:
    Jeffrey T. Leek
  • 依托单位:
A massive study of data science to address the scientific reproducibility crisis
  • 批准号:
    9100338
  • 项目类别:
  • 资助金额:
    $36.45万
  • 财政年份:
    2016
  • 负责人:
    Jeffrey T. Leek
  • 依托单位:
海外基金