课题基金 / 基金详情

Category II: Unlocking Interactive AI Development for Rapidly Evolving Research

Category II: Unlocking Interactive AI Development for Rapidly Evolving Research
第二类:为快速发展的研究解锁交互式人工智能开发
批准号:
2005597
负责人:
Paola Buitrago
金额:
$500.0万
依托单位:
依托单位国家:
美国
项目类别:
Cooperative Agreement
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-06-01 至 2025-05-31

项目摘要

项目成果

Paola Buitrago的其他基金

相似基金

相关文献

中文摘要
翻译
高精度的人工智能(AI)一再在科学和工程领域带来巨大的利益。人工智能能够对大型数据集进行信息提取和分析,并创建高保真模型,这些模型可以增强或取代传统仿真代码中计算成本更高的计算。通过这种方式,人工智能可以将科学研究的时间加快几个数量级。为了获得这些好处,必须首先训练高度复杂的人工智能模型。训练是一个计算密集型的过程,需要几天、几周甚至几个月的时间,需要对网络架构及其超参数进行优化,并且限制了所处理挑战的范围和复杂性。如果训练时间可以减少到互动的程度,即使在更极端的情况下也只需要几分钟或几个小时,那会怎么样?结果将是变革性的:科学家和工程师可以迅速发展和完善他们的想法,使他们能够为最紧迫和最复杂的问题提供高影响力的解决方案。为了通过实现前所未有的人工智能速度和可扩展性来帮助推进知识,卡内基梅隆大学和匹兹堡大学的联合研究中心匹兹堡超级计算中心(PSC)将与Cerebras Systems和惠普企业(HPE)合作,部署Neocortex,这是一种创新的计算资源,将通过大大缩短深度学习训练/推理所需的时间来加速科学发现。促进深度人工智能模型与科学工作流程的更大整合,并为开发更高效的人工智能和图形分析算法提供革命性的创新硬件。Neocortex将通过加速科学研究、开发更精确的模型和使用更大的训练数据、将模型并行性扩展到前所未有的水平、通过简化调优和超参数优化专注于人类生产力,以及为探索新领域提供革命性的硬件平台来推进知识。Neocortex将为NSF网络基础设施生态系统引入最强大的人工智能处理器,并将使改变游戏规则的计算能力民主化,否则只有科技巨头才能获得,对于学生、博士后、教师和其他需要更快的培训周转来分析数据并将人工智能与模拟集成的人来说。它将提供一个独特的机会,探索突破性的新人工智能硬件架构的潜力,利用Cerebras CS-1人工智能平台的革命性人工智能处理器技术和HPE Superdome Flex的大型内存扩展功能,解锁新的见解并加快发现的时间。Neocortex项目还将侧重于围绕这些革命性能力建立一个强大的社区,包括与其他领先的国家机构合作,并强调包容性和多样性。它将通过培训和实习培养STEM人才,通过产业拓展发展美国劳动力和国家竞争力,并促进国际合作。公共宣传和XSEDE校园冠军和领域冠军活动将有助于吸引更广泛的受众。新颖的Neocortex架构将两个异常强大的Cerebras CS-1 AI服务器与一个超大共享内存的HPE Superdome Flex HPC服务器相结合,以实现前所未有的AI可扩展性和出色的系统平衡。每台Cerebras CS-1都由一个Cerebras晶圆规模引擎(WSE)处理器驱动,这是一种革命性的高性能处理器,专门用于加速深度学习训练和推理。Cerebras WSE是迄今为止最大的芯片,包含40万个人工智能优化核心,在46225平方毫米的晶圆上实现,拥有1.2万亿个晶体管。片上结构通过完全可配置的2D网格提供100Pb/s的带宽,没有软件开销。Cerebras WSE包括18GB的SRAM,可在一个时钟周期内以9PB/s的带宽访问。Cerebras WSE经过独特的设计,可以实现高效的稀疏计算,既不浪费时间,也不浪费功率,乘以深度网络中的许多零。Cerebras CS-1软件可以使用常见的ML框架(如TensorFlow和PyTorch)进行编程,为了提高计算效率,这些框架被映射到优化的图形表示和一组特定于模型的计算内核上。它还支持本地代码开发。支持最流行的深度学习框架和自动、透明的加速将为研究人员提供卓越的易用性。Neocortex的HPE Superdome Flex HPC服务器将是Cerebras CS-1服务器的一个非常强大、用户友好的前端。这将使进出附加WSE的数据进行灵活的预处理和后处理,防止瓶颈,充分利用WSE的能力,并实现高级深度学习功能,如增强、超参数和模型优化以及集成学习。Superdome Flex将配备24TB RAM、204.8TB高性能NVMe闪存、32个Intel Xeon cpu和24个100GbE网络接口卡,为跨多个CS-1系统扩展应用程序创造最大的灵活性。在内部,HPE Superdome Flex由一个定制的存储结构ASIC连接,用于缓存一致的硬件共享内存,保持850GB/s的互连带宽。它的大而快的内存和高计算性能将使非常大的数据集训练非常容易,避免了拆分和尝试跨工作节点负载平衡数据集的繁重任务。每个Cerebras CS-1具有1.2Tbps的I/O,并将通过12条标准100GbE链路连接到HPE Superdome Flex。这种配置将提供最大的性能和灵活性,包括通过PSC-Cerebras和PSC-HPE的研究合作伙伴关系,探索将训练扩展到多个CS-1系统。Neocortex将通过16个InfiniBand HDR100连接(总带宽1.6Tbps)与nsf支持的容量资源Bridges-2进行联合。这种联合将为用户社区带来巨大的好处,包括访问Bridges-2文件系统来管理持久数据;数据预处理和传统机器学习的通用计算;使用Bridges-2与数据密集型项目进行互操作;以及与其他XSEDE服务提供商、校园、实验室和云的高带宽外部网络连接。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
High-accuracy artificial intelligence (AI) has repeatedly delivered great benefit across science and engineering. AI enables information extraction and analysis of large datasets and the creation of high-fidelity models that can augment or replace more computationally expensive calculations in traditional simulation codes. In this way, AI can accelerate time-to-science by orders of magnitude. To reap these benefits, highly complex AI models must first be trained. Training is a computation-intensive process that takes days, weeks, or even months, requires optimization of both the network architecture and its hyperparameters, and limits the scope and complexity of challenges addressed. What if training time could be reduced to the point of being interactive, taking only minutes or hours even in the more extreme cases? The result would be transformative: scientists and engineers could rapidly develop and refine their ideas, enabling them to achieve high-impact solutions to the most pressing and complex issues.To help advance knowledge by enabling unprecedented AI speed and scalability, the Pittsburgh Supercomputing Center (PSC), a joint research center of Carnegie Mellon University and the University of Pittsburgh, in partnership with Cerebras Systems and Hewlett Packard Enterprise (HPE), will deploy Neocortex, an innovative computing resource that will accelerate scientific discovery by vastly shortening the time required for deep learning training/inference, foster greater integration of deep AI models with scientific workflows, and provide revolutionary innovative hardware for the development of more efficient algorithms for artificial intelligence and graph analytics. Neocortex will advance knowledge by accelerating scientific research, enabling development of more accurate models and use of larger training data, scaling model parallelism to unprecedented levels, focusing on human productivity by simplifying tuning and hyperparameter optimization, and providing a revolutionary hardware platform for the exploration of new frontiers.Neocortex will introduce the most powerful AI processor to the NSF cyberinfrastructure ecosystem and will democratize access to game-changing compute power, otherwise only available to tech giants, for students, postdocs, faculty, and others, who require faster training turnaround to analyze data and integrate AI with simulation. It will provide a unique opportunity to explore the potential of a groundbreaking new AI hardware architecture, tapping into the revolutionary AI processor technology of the Cerebras CS-1 AI platform and the large in-memory scale up capabilities of HPE Superdome Flex to unlock new insights and accelerate time to discovery. The Neocortex project will additionally focus on building a strong community around these revolutionary capabilities, including collaborations with other leading national institutions and emphasizing inclusion and diversity. It will build STEM talent through training and internships, develop the U.S. workforce and national competitiveness through industrial outreach, and foster international collaborations. Public outreach and XSEDE campus champion and domain champion activities will help engage a wider audience.The novel Neocortex architecture will couple two exceptionally powerful Cerebras CS-1 AI servers with an exceptionally large shared memory HPE Superdome Flex HPC server to achieve unprecedented AI scalability with excellent system balance. Each Cerebras CS-1 is powered by one Cerebras Wafer Scale Engine (WSE) processor, a revolutionary high-performance processor designed specifically to accelerate deep learning training and inferencing. The Cerebras WSE is the largest chip ever built, containing 400,000 AI-optimized cores implemented on a 46,225 square millimeter wafer with 1.2 trillion transistors. An on-chip fabric provides 100Pb/s of bandwidth through a fully configurable 2D mesh with no software overhead. The Cerebras WSE includes 18GB of SRAM accessible within a single clock cycle at 9PB/s bandwidth. The Cerebras WSE is uniquely engineered to enable efficient sparse computation, wasting neither time nor power multiplying the many zeroes that occur in deep networks. The Cerebras CS-1 software can be programmed with common ML frameworks such as TensorFlow and PyTorch, which for computational efficiency are mapped onto an optimized graph representation and a set of model-specific computation kernels. It also supports native code development. Support for the most popular deep learning frameworks and automatic, transparent acceleration will provide researchers with exceptional ease of use.The HPE Superdome Flex HPC server of Neocortex will be an extremely powerful, user-friendly front end for the Cerebras CS-1 servers. This will enable flexible pre- and post-processing of data flowing in and out of the attached WSEs, preventing bottlenecks to taking full advantage of the WSE capability, and implementing advanced deep learning functions such as augmentation, hyper-parameter and model optimization, and ensemble learning. The Superdome Flex will be robustly provisioned with 24TB of RAM, 204.8TB of high-performance NVMe flash storage, 32 Intel Xeon CPUs, and 24 100GbE network interface cards to create the greatest flexibility for scaling applications across multiple CS-1 systems. Internally, the HPE Superdome Flex is interconnected by a custom memory fabric ASIC for cache-coherent hardware shared memory sustaining 850GB/s of interconnect bandwidth. Its large and fast memory and high compute performance will enable training on very large datasets with exceptional ease, avoiding the laborious task of splitting and trying to load-balance datasets across worker nodes.Each Cerebras CS-1 has 1.2Tbps I/O, and will connect to the HPE Superdome Flex via twelve, standard 100GbE links. This configuration will deliver the greatest possible performance and flexibility, including, via PSC-Cerebras and PSC-HPE research partnerships, exploration of scaling training to multiple CS-1 systems.Neocortex will be federated via 16 InfiniBand HDR100 connections (an aggregate 1.6Tbps) with Bridges-2, an NSF-supported capacity resource. This federation will yield great benefits to the user community including access to the Bridges-2 filesystem to manage persistent data; general-purpose computing for data preprocessing and traditional machine learning; interoperation with data-intensive projects using Bridges-2; and high-bandwidth external network connectivity to other XSEDE Service Providers, campus, labs, and clouds.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Deep Learning Benchmark Studies on an Advanced AI Engineering Testbed from the Open Compass Project
开放罗盘项目的高级人工智能工程测试平台的深度学习基准研究
DOI: 10.1145/3569951.3597596
发表时间: 2023
期刊: PEARC '23: Practice and Experience in Advanced Research Computing
影响因子: --
作者: [Wang, Mei-Yu, Uran, Julian, Buitrago, Paola]
通讯作者: Buitrago, Paola
Collaborative Research: CyberTraining: Pilot: Building a strong community of computational researchers empowered in the use of novel cutting-edge technologies
  • 批准号:
    2320991
  • 项目类别:
    Standard Grant
  • 资助金额:
    $6.18万
  • 财政年份:
    2023
  • 负责人:
    Paola Buitrago
  • 依托单位:
Open Compass: Leveraging the Compass AI Engineering Testbed to Accelerate Open Research
  • 批准号:
    1833317
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2018
  • 负责人:
    Paola Buitrago
  • 依托单位:
国内基金
海外基金
基于生境成像与深度学习联合临床特征构建II型卵巢癌术前淋巴结转移预测模型的研究
鸡软骨非变性II型胶原高效制备和靶向递送的关键技术开发与应用示范
青蒿琥酯协同TROP2/线粒体级联靶向的NIR-II多模态诊疗用于晚期TNBC精准诊断与治疗的机制研究
  • 批准号:
    2026JJ30126
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    杨沙
  • 依托单位:
苏合颗粒治疗慢性萎缩性胃炎的临床(II期)评价关键技术研究