课题基金 / 基金详情

NSCI Elements: Software - PFSTRASE - A Parallel FileSystem TRacing and Analysis SErvice to Enhance Cyberinfrastructure Performance and Reliability

NSCI Elements: Software - PFSTRASE - A Parallel FileSystem TRacing and Analysis SErvice to Enhance Cyberinfrastructure Performance and Reliability
NSCI Elements:软件 - PFSTRASE - 用于增强网络基础设施性能和可靠性的并行文件系统跟踪和分析服务
批准号:
1835135
负责人:
Richard Evans
金额:
$38.59万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-10-01 至 2022-09-30

项目摘要

项目成果

Richard Evans的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将开发一种开源软件服务,即并行文件系统跟踪和分析服务(PFSTRASE),以提高国家数据存储系统的可靠性和性能。美国最大的超级计算机。由于模拟和计算更忠实地代表了现实,因此它们的规模随着它们消耗和生成的数据的大小而相应增长。为了处理这些数据的存储和移动,超级计算系统建立在大规模并行数据存储系统的主干上。由于它们的并行特性,这些存储系统能够以传统存储系统数百倍的速度移动数据,从而实现不切实际的计算。这些存储系统提供的性能能力伴随着复杂性,这导致它们的功能通常远远低于最佳状态,甚至在某些情况下会失败。这导致浪费了计算时间,最终失去了科学进步。对于当前和未来的计算系统来说,能够解决这些问题并提高存储系统可靠性和性能的工具的开发状态是不够的。PFSTRASE将通过持续和自动监控存储系统的健康和性能来填补这一空白,通过易于使用的界面提供见解,这将提高存储和超级计算机系统的可靠性和性能。并行文件系统(pfs)是高性能计算(HPC)体系结构中最关键的高可用性组件,为运行的计算、用户和系统服务运行的环境以及应用程序和数据的存储提供输入/输出(I/O)服务。由于这个中心角色,PFS中的故障或性能下降事件会影响HPC资源的每个用户。系统管理员必须快速有效地处理PFS事件;但是,通常没有足够的资料来确定PFS活动和事件之间的确切因果关系,妨碍了及时和有针对性的补救措施的执行。为了填补这一信息缺口,将开发一个开源的并行文件系统跟踪和分析服务(PFSTRASE),该服务跟踪和分析必要的数据,以建立PFS活动与已实现和即将发生的事件之间的因果关系。该项目将为开源Lustre文件系统实现服务,这是大规模HPC站点中最常用的PFS。将测量特定PFS目录和文件操作的负载,并将其合并到服务中,以构建来自每个作业、进程和用户的真实服务器负载贡献。服务吗?s的基础设施将持续监控整个PFS,并生成实时、无缝的表示,将作业、进程和用户的贡献与存储服务器负载、网络带宽和存储容量连接起来。该基础设施将提供一个易于导航的web界面,以可视化的格式呈现实时和历史数据。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project will develop an open-source software service, the Parallel FileSystem TRacing and Analysis SErvice (PFSTRASE), that improves the reliability and performance of data storage systems for the nation?s largest supercomputers. As simulations and computations represent reality more faithfully they grow commensurately in scale along with the size of the data they consume and generate. To handle the storage and movement of this data, supercomputing systems are built on the backbone of massively parallel data storage systems. Due to their parallel nature these storage systems are capable of moving data at hundreds of times the speed of conventional storage systems, enabling otherwise impractical computations. The performance capabilities these storage systems provide is accompanied by a complexity that results in them often functioning significantly less than optimally and even in some instances failing. This results in wasted computational time and ultimately lost scientific progress. The state of development of tools that could cast light on these problems and improve storage system reliability and performance is inadequate for current and future computing systems. PFSTRASE will fill this gap by continually and automatically monitoring storage system health and performance, providing insights through an easy to use interface that will improve the reliability and performance of storage and supercomputer systems. Parallel filesystems (PFSs) are the most critical high-availability components of High Performance Computing (HPC) architectures, providing input/output (I/O) services to running computations, the environment that users and system services operate in, and storage for applications and data. Because of this central role, failure or performance degradation events in the PFS impact every user of an HPC resource. PFS events must be dealt with quickly and effectively by system administrators; however, there is typically insufficient information to establish precise causal relationships between PFS activity and events, impeding the implementation of timely and targeted remedies. To fill this information gap, an open-source Parallel FileSystem TRacing and Analysis SErvice (PFSTRASE) that traces and analyzes the requisite data to establish causal relationships between PFS activity and both realized and imminent events will be developed. This project will implement the service for the open-source Lustre filesystem, which is the most commonly used PFS at large-scale HPC sites. Loads for specific PFS directory and file operations will be measured and incorporated into the service to construct authentic server load contributions from every job, process, and user. The service?s infrastructure will continuously monitor the entire PFS and generate a real-time, seamless representation that connects contributions of jobs, processes, and users to storage server loads, network bandwidth, and storage capacities. The infrastructure will provide an easily navigable web interface that presents this data, both real-time and historical, in a visual format.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Democratizing Parallel Filesystem Monitoring
并行文件系统监控民主化
DOI: 10.1109/cluster49012.2020.00065
发表时间: 2020
期刊: 2020 IEEE International Conference on Cluster Computing (CLUSTER
影响因子: --
作者: [Evans, Richard Todd]
通讯作者: Evans, Richard Todd
COLLABORATIVE RESEARCH: We are thriving: Challenging negative discourse through voices of women in project teams
  • 批准号:
    2015741
  • 项目类别:
    Standard Grant
  • 资助金额:
    $11.8万
  • 财政年份:
    2020
  • 负责人:
    Richard Evans
  • 依托单位:
Size, shape and surface properties in realistic models of magnetic nanocrystals
  • 批准号:
    EP/P022006/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $12.82万
  • 财政年份:
    2017
  • 负责人:
    Richard Evans
  • 依托单位:
Mapping "missing" conformations of ATP-gated P2X receptor ion channels
  • 批准号:
    BB/P001076/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $53.96万
  • 财政年份:
    2016
  • 负责人:
    Richard Evans
  • 依托单位:
Cross-linking and molecular modelling to determine the structure and dynamics of the intracellular regions of ATP gated P2X receptor ion channels
  • 批准号:
    BB/M000990/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $44.03万
  • 财政年份:
    2014
  • 负责人:
    Richard Evans
  • 依托单位:
海外基金