EAGER: Bridging the Last Mile; Towards an Assistive Cyberinfrastructure for Accelerating Computationally Driven Science
EAGER: Bridging the Last Mile; Towards an Assistive Cyberinfrastructure for Accelerating Computationally Driven Science
批准号:
1945347
负责人:
Rajiv Ramnath
金额:
$29.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-01-01 至 2024-06-30
中文摘要
尽管计算科学家拥有高速计算机和强大的软件工具,但他们进行科学发现的能力受到了限制,因为他们的大部分时间都被低效地花在数据探索、操纵和可视化、初步分析以及偶然尝试为他们的软件工具找到正确的设置以有效使用,而不是大规模运行他们的计算上。例如,进行基因测序(识别生物体DNA中的单个碱基核苷酸的过程)的研究人员必须计算出使用哪种测序方法,使用几个软件中的哪个软件包,使用什么设置的软件包,等等。研究人员将从先前使用他们所掌握的软件工具中学习,即从他们自己和该领域其他研究人员过去的经验和错误中学习。然而,只有一小部分人能够做到这一点。这一探索性项目将调查是否可以从过去使用人工智能技术的具体应用经验中提取有效的智能指导,然后以量身定制的方式提供给个别研究人员。这将在研究人员正在使用的软件工具中完成,并由基础网络基础设施本身提供,因此任何使用这些工具的人都可以受益。该项目将探索提供此类援助的基于人工智能的方法的可行性,并对一系列能力进行原型、试点和验证,并广泛传播结果。目前,科学网络基础设施(CI)的研究和开发侧重于扩展CI:开发可更好地扩展的算法,自动化管道,构建可扩展的计算和数据系统,以及优化系统资源分配。另一方面,该项目侧重于衡量研究人员的个体,即通过应用一系列技术使她更有效率,从科学家-CI交互作用的观察性研究,到端到端仪器,用于记录和跟踪跨CI的这些交互作用,从这些交互作用建立机器学习的模型,以及在人机合作范例中嵌入和试验这些模型。通过将这些技术应用于侧重于具体应用的方法和技术框架,这些技术将获得最好的成功机会。建议的方法将改变参与者的专业知识,测试绩效的假设,并使用透明?引导我?在设计指南时确定个体差异维度的方法。样本领域的科学是基因组学,其目标是对选定物种的完整DNA进行排序和组装,这一项目的工作可能特别具有变革性。项目团队是综合性、跨学科和趋同的;研究人员拥有基因组学、软件工程、系统、数据科学、项目管理和人机系统方面的专业知识,正在与俄亥俄州超级计算机中心的一家关键CI提供商合作,并集体为一小群研究生提供建议。该项目与NSF关于利用数据革命、不断增长的融合研究和工作的未来以及软件网络基础设施高级基础设施标准办公室的大想法保持一致,因为它是领域和计算机科学驱动的、创新的、协作和融合的、战略管理的,并建立在NSF先前的重大投资基础上的-这是一条通向可持续发展的明确道路。这项工作可能会使来自许多科学领域的计算科学家变得更有生产力,从而加速发现。其他更广泛的影响是通过计算科学的教育案例研究、对仪器标准的贡献以及对传播与信息的观察性和经验性研究方法。这项工作的另一个贡献是评估CI工具的可用性。这将直接使工具设计人员能够构建更有用的工具。强调扩大参与:主要调查员之一是一名妇女。这三个私人助理都有招募女学生和与女学生一起工作的历史。至少有一名受资助的研究生将是即将入学的女学生。此外,该项目团队还将与基因组实验室的两名女学生和一名女性博士后研究员密切合作。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Even though computational scientists have high-speed computers and powerful software tools at their disposal, their ability to make scientific discovery is held back because much of their time is inefficiently spent in data exploration, manipulation and visualization, preliminary analyses, and hit-and-miss attempts to find the right settings for their software tools to be used effectively, rather than in running their computations at scale. For example, researchers doing gene sequencing (a process in which the individual base nucleotides in an organism's DNA are identified) have to figure what sequencing method to use, which software package - out of several - to use for that sequencing method, what settings to use for the software package, and so on. Researchers would greatly benefit by learning from prior use of the software tools at their disposal, that is, from past experiences and mistakes, both their own, and those of other researchers in the field. However, only a small percentage of them are able to do so. This exploratory project will investigate whether effective, intelligent guidance can be extracted from past experience on specific applications using artificial intelligence techniques and then provided in a tailored manner to the individual researcher. This will be done within the software tools being used by the researcher, and provided by the underlying cyberinfrastructure itself, so that anyone who uses these tools can benefit. This project will explore the feasibility of AI-based approaches for providing such assistance, and prototype, pilot and validate a range of capabilities, and widely disseminate results.Current research and development in cyberinfrastructure(CI) for science focuses on scaling the CI: developing algorithms that scale better, automating pipelines, building scalable computational and data systems, and optimizing system resource allocation. This project, on the other hand, focuses on scaling the individual researcher, i.e. making her more effective, through the application of a portfolio of techniques, from observational studies of scientist-CI interactions, to end-to-end instrumentation for recording and tracking these interactions across the CI, building machine-learned models from these interactions, and embedding and experimenting with these models within human-machine teaming paradigms. These techniques will be given the best chance to succeed by applying them in a methodological and technical framework that focuses on specific applications. The proposed methods will vary participant expertise, test hypotheses of performance, and use transparent ?guide me? methodologies to establish dimensions of individual differences in designing guidance. The exemplar domain science is genomics, where the goal is to sequence and assemble the complete DNA of selected species, and where the work in this project could be particularly transformative. The project team is integrative, interdisciplinary and convergent; the investigators have expertise in genomics, software engineering, systems, data science, project management, and human-machine systems, and are working with a key CI provider in the Ohio Supercomputer Center, and collectively advising a small team of graduate students. The project is aligned with the NSF Big Ideas of Harnessing the Data Revolution, Growing Convergence Research and the Future of Work, and the Office of Advanced Infrastructure criteria for software cyberinfrastructure, since it is domain and computer sciences-driven, innovative, collaborative and convergent, strategically managed, and building on significant prior investments by NSF - a clear path to sustainability. This work could make computational scientists from many science domains transformatively more productive, leading to accelerated discovery. Additional broader impacts are through educational case-studies for computational science, contributions to instrumentation standards, and observational and empirical study methods for CI. A side contribution of this work will be in assessing the usability of CI tools . This will directly enable tool designers to build more usable tools. Broadening participation has been emphasized: one of the principal investigators is a woman. All three PIs have a history of recruiting and working with women students. At least one of the funded graduate students will be an incoming female student. The project team will additionally be closely collaborating with two women students and a woman post-doctoral researcher in the genomics laboratory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Establishing a Generalizable Framework for Generating Cost-Aware Training Data and Building Unique Context-Aware Walltime Prediction Regression Models
建立通用框架来生成成本感知训练数据并构建独特的上下文感知 Walltime 预测回归模型
DOI:
--
发表时间:
2022
期刊:
20th IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA 2022
影响因子:
--
作者:
[Swathi Vallabhajosyula, Rajiv Ramnath]
通讯作者:
Rajiv Ramnath
Towards Practical, Generalizable Machine-Learning Training Pipelines to build Regression Models for Predicting Application Resource Needs on HPC Systems
迈向实用、可推广的机器学习培训管道,构建回归模型来预测 HPC 系统上的应用程序资源需求
DOI:
--
发表时间:
2022
期刊:
Practice and Experience in Advanced Research Computing (PEARC
影响因子:
--
作者:
[Swathi Vallabhajosyula, Rajiv Ramnath]
通讯作者:
Rajiv Ramnath
SHF: Small: Techniques and Frameworks for Exploiting Recent SIMD Architectural Advances
-
批准号:1526386
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2015
-
负责人:Rajiv Ramnath
-
依托单位:
Curriculum for Accelerated Services Engineering (CASE)
-
批准号:0837555
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2009
-
负责人:Rajiv Ramnath
-
依托单位:
海外基金