课题基金 / 基金详情

CyberTraining: Implementation: Small: Developing a Best Practices Training Program in Cyberinfrastructure-Enabled Machine Learning Research

CyberTraining: Implementation: Small: Developing a Best Practices Training Program in Cyberinfrastructure-Enabled Machine Learning Research
网络培训:实施:小型:制定网络基础设施支持的机器学习研究最佳实践培训计划
批准号:
2017767
负责人:
Mary Thomas
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2024-08-31

项目摘要

项目成果

Mary Thomas的其他基金

相似基金

相关文献

中文摘要
翻译
该项目面向NSF研究人员,他们需要开发在国家网络基础设施(CI)上运行的机器学习应用程序。该项目的核心是支持CI的机器学习(CIML)培训系统和存储库,它将支持ML空间中网络素养的发展。与目前在网上很容易找到的许多ML相关培训材料不同,CIML材料将集中在使用CI支持的ML技术的科学和工程应用程序上,并将用于培训一支能够理解使用CI、新HPC架构、软件和应用程序的挑战的研究队伍。CIML培训材料的用户社区将包括学生(本科生、研究生)、博士后、PI、研究人员、教育工作者和HPC培训人员,每个人都有自己的不同背景和应用要求。该项目将支持国家教育目标,确保传播与信息单元在先进的传播与信息工具和资源上运行,并将先进传播与传播的核心扫盲和学科适当技能纳入课程和教学材料。CIML将通过促进一支能够在气候和天气、生物科学、物理和化学等科学领域开发ML应用的劳动力来支持国家安全问题。作为外展和扩展培训努力的结果,该计划将影响数千名用户,并帮助培养下一代CI研究人员。CIML培训材料将在网上提供,因此该项目具有巨大的潜力,可以超越NSF的网络工作人员,影响其他社区,包括医院和医疗系统、交通和电力监控系统、股市监控系统和灾难应对系统。支持网络基础设施的机器学习(CIML)培训系统和资源库将使用“最佳实践”方法来开发一个独特的计划,目标是将机器学习(ML)和大数据分析方法用于其领域特定应用程序或大规模网络基础设施上的教学材料的研究人员。该项目将应用网络扫盲和高性能计算能力的方法,根据学习的各个方面,从侧重技术到注重解决问题或侧重于ML或计算科学,确定一套核心的ML和特定领域的扫盲领域。CIML系统的来源将来自高性能计算培训、现有的高性能计算研究人员和用户、合作者以及新代码和方法的工作。开发的材料将通过CIML存储库获得,其中包括一个网站、文档、GitHub代码、数据和相关材料的存储库。Cimil将成为两个社区的有用工具:想要了解他们需要掌握什么技术和技能才能运行特定ML应用程序的用户,使用什么系统,以及建议的软件库;以及需要知道要教授什么主题的培训人员。这些努力的结果将产生一个由机器学习和数据分析、CI用户(CIU)和贡献者(CIC)组成的社区,他们积极为培训材料储存库做出贡献,并将材料纳入他们的项目和课程。作为这些努力的结果,CIML计划将通过开发基于网络基础设施的材料来扩大整个研究人员正在进行的教育和培训的范围,这些材料将利用并促进为XSEDE培训、高等教育和其他计划开发的培训材料,并将影响数以千计的现有和新用户,包括学生(本科生/研究生)、博士后、PI、研究人员和教育工作者,每个人都有自己不同的背景和应用要求。CIML培训材料将在网上提供,因此该项目具有巨大的潜力,可以覆盖到网络劳动力之外,并影响许多社区,包括医院和医疗系统、交通和电力监控系统、股票和市场监控系统以及灾害响应系统。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project is targeted towards the NSF research workforce who need to develop machine learning applications that will run on the national cyberinfrastructure (CI). At the heart of this project is the CI-enabled Machine Learning (CIML) training system and repository that will support the development of cyberliteracy in the ML space. Unlike much of current ML-related training material found easily online, CIML material will be centered on science and engineering applications that make use of CI-enabled ML techniques and will be used to train a research workforce that is capable of understanding the challenges of working with CI, new HPC architectures, software, and applications. The user community for CIML training material will include students (undergraduate, graduate), postdocs, PIs, researchers, educators, and HPC trainers, each with their own diverse backgrounds and application requirements. The project will support national educational goals by ensuring that the CI modules run on advanced CI tools and resources, and that core literacy and discipline appropriate skills in advanced CI will be integrated into curricula and instructional material. CIML will support national security concerns by facilitating a workforce capable of developing ML applications in scientific domains such as climate and weather, the biosciences, physics, and chemistry. As a result of outreach and extension of the training efforts, this program will impact thousands of users and help develop the next generation of the CI research workforce. CIML training material will be available online, so the project has a huge potential to reach beyond the NSF cyber workforce to impact other communities including hospital and medical treatment systems, transportation and electrical monitoring systems, stock market monitoring systems, and disaster response systems. The Cyberinfrastructure-enabled Machine Learning (CIML) training system and repository will use a “best practices” approach to develop a unique program targeted towards the research workforce who use machine learning (ML) and big data analytics methods for their domain specific applications or instructional material on large-scale cyberinfrastructure. The project will apply methods of Cyber Literacy and HPC Competencies to define a set of core ML and domain specific literacy areas as a function of the dimensions of learning ranging from a technological focus to a problem-solving focus or a focus on ML or computational science. Sources for the CIML system will be drawn from the work of HPC training, existing HPC researchers and users, collaborators, as well as new code and methods. The materials developed will be available via the CIML repository, which includes a web site, documentation, GitHub repositories for code, data, and related materials. CIMIL will become a useful tool for 2 communities: users who want to understand what technologies and skills they need to master in order to run a particular ML application, what systems to use, and suggested software libraries; and trainers who need to know what topics to teach. The outcome of these efforts will result in a community of machine learning and data analytics CI Users (CIU) and Contributors (CIC) who actively contribute to the training material repository and incorporate the materials into their projects and courses. As a result of these efforts, the CIML program will extend the scope of the ongoing education and training across the research workforce by developing cyberinfrastructure-based materials that will utilize and contribute to training material developed for XSEDE training, higher education, and other programs, and will impact thousands of existing and new users, including students (undergrads/grads), postdocs, PIs, researchers, and educators, each with their own diverse backgrounds and application requirements. CIML training material will be available online, so the project has a huge potential to reach beyond the cyber workforce and to impact many communities, including hospital and medical treatment systems, transportation and electrical monitoring systems, stock and market monitoring systems, and disaster response systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Experiences in Building a User Portal for Expanse Supercomputer
Expanse超级计算机用户门户构建经验
DOI: 10.1145/3437359.3465590
发表时间: 2021
期刊: Practice & Experience in Advanced Research Computing (PEARC
影响因子: --
作者: [Sakai, Scott, Mishin, Dmitry, Sivagnanam, Subhashini, Tatineni, Mahidhar, Kandes, Martin, Thomas, Mary, Irving, Christopher, Strande, Shawn, Norman, Michael]
通讯作者: Norman, Michael
Critique of: “A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery” by SCC Team From UC San Diego
加州大学圣地亚哥分校 SCC 团队对“通过马尔可夫毯子发现进行基于约束的贝叶斯网络学习的并行框架”的评论
DOI: 10.1109/tpds.2022.3217284
发表时间: 2023
期刊: IEEE Transactions on Parallel and Distributed Systems
影响因子: 5.3
作者: [Gupta, Arunav, Ge, John, Li, John, Kong, Zihao, He, Kaiwen, Mikhailov, Matthew, Chin, Bryan, Li, Xiaochen, Apodaca, Max, Rodriguez, Paul]
通讯作者: Rodriguez, Paul
Expanse : Computing without Boundaries
Expanse™:无边界计算
DOI: --
发表时间: 2021
期刊: Practice & Experience in Advanced Research Computing (PEARC
影响因子: --
作者: [Strande, Shawn, Altintas, Ilkay, Cai, Haisong, Cooper, Trevor, Irving, Christopher, Kandes, Marty, Majumdar, Amitava, Mishin, Dmitry, Perez, Ismael, Pfeiffer, Wayne]
通讯作者: Pfeiffer, Wayne
CyberTraining: CIP: Training and Developing a Research Computing and Data CI Professionals (RCD-CIP) Community
  • 批准号:
    2230127
  • 项目类别:
    Standard Grant
  • 资助金额:
    $670.2万
  • 财政年份:
    2022
  • 负责人:
    Mary Thomas
  • 依托单位:
NMI: Collaborative Proposal: Middleware for Grid Portal Development
NMI: Collaborative Proposal: Middleware for Grid Portal Development
  • 批准号:
    0330652
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $58.78万
  • 财政年份:
    2003
  • 负责人:
    Mary Thomas
  • 依托单位:
海外基金