XPS:FULL:DSD: Collaborative Research: FPGA Cloud Platform for Deep Learning, Applications in Computer Vision
XPS:FULL:DSD: Collaborative Research: FPGA Cloud Platform for Deep Learning, Applications in Computer Vision
批准号:
1533739
负责人:
Michael Ferdman
金额:
$57.4万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2020-08-31
中文摘要
我们即将在深度学习应用方面取得巨大进步,这将很快使基于计算机视觉的识别在科学研究、商业应用和日常生活中得到实用和广泛采用。巨大的挑战问题就在我们的触手可及之处;我们很快就能够建立能够识别我们看到的几乎所有东西的自动化系统,能够识别心理学家假设的人类可以识别的数万个基本级别类别的系统,能够不断从照片、视频和网络内容中学习的系统,以便创建更完整、更准确的世界视觉模型。然而,尽管深度学习的计算能力显然触手可及,但同样明显的是,所需的计算能力不可能来自通用处理器。为了取得成功,我们将需要基于硬件加速器构建专门的领域特定计算系统,这些系统能够利用深度学习工作负载固有的极细粒度并行。该项目利用并行化和可重新配置的硬件来创建一个自动化系统,将计算机视觉算法分布到大量现场可编程门阵列(现场可编程门阵列)上。该项目建立在领域特定硬件生成工具的最新进展的基础上,以便将潜在的并行性和性能功耗比优势带到大规模计算机视觉问题中。通过开发一个在大规模现场可编程门阵列的云上运行深度学习算法的平台,该提议明确地解决了扩展算法超出单个芯片所能处理的范围。这包括解决算法分析中的一系列具有挑战性的问题,构建特定于领域的硬件生成器,跨多个FGA扩展算法的通信,以及为应用于计算机视觉问题的最先进的深度学习方法生成硬件的广泛验证。该项目推进了用于设计特定于领域的算法的FPGA实现的工具,朝着使具有更大并行性的更高效的计算更广泛地可用迈出了一步。特别是,对于计算机视觉来说,多项改进的产品将带来显著的好处:更高的并行度,尽可能移动到固定点的更低的门要求,以及更好的性能功耗比,从而在服务器中产生更高的计算密度。总而言之,这些都有可能显著提高计算机视觉在我们日常生活中的作用,使计算机能够更好地理解我们世界的背景。
英文摘要
We stand on the verge of dramatic advances in deep learning applications, which will soon enable practicality and widespread adoption of computer vision based recognition in scientific inquiry, commercial applications, and everyday life. Grand challenge problems are within our reach; we will soon be able to build automated systems that recognize nearly everything we see, systems that can recognize the tens of thousands of basic-level categories that psychologists posit humans can recognize, systems that continuously learn from photos, video, and web content in order to create more complete and accurate visual models of the world. However, while it is clear that the computational capabilities for deep learning are within reach, it is equally clear that the required computational power cannot come from general-purpose processors. To succeed, we will need to build specialized domain-specific computing systems based on hardware accelerators that are capable of exploiting the extreme fine-grained parallelism inherent in deep-learning workloads. This project leverages parallelization and reconfigurable hardware to create an automated system that distributes computer vision algorithms onto a large number of field-programmable gate arrays (FPGA Cloud). This project builds on recent advances in domain-specific hardware generation tools in order to bring the potential parallelism and performance per watt advantages of FPGAs to large-scale computer vision problems. By developing a platform to run deep learning algorithms on large clouds of FPGAs, this proposal explicitly addresses scaling algorithms beyond what a single chip can process. This involves addressing a wide range of challenging problems in algorithm analysis, building domain-specific hardware generators, communication for scaling algorithms across multiple FPGAs, and extensive validation of generating hardware for state-of-the-art deep learning approaches applied to computer vision problems. This project advances tools for designing domain-specific FPGA implementations of algorithms, taking a step toward making more efficient computing with greater parallelism more widely available. In particular, for computer vision, there will be significant benefits from a product of multiple improvements: higher parallelism, lower gate requirement by moving to fixed point when possible, and better performance per watt leading to higher computation density in servers. Together, these have the potential to significantly increase the extent to which computer vision can be a part of our daily lives, making computers better able to understand the context of our world.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
A Full-System VM-HDL Co-Simulation Framework for Servers with PCIe-Connected FPGAs
适用于具有 PCIe 连接 FPGA 的服务器的全系统 VM-HDL 联合仿真框架
DOI:
10.1145/3174243.3174269
发表时间:
2018
期刊:
FPGA '18 Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
作者:
[Cho, Shenghsun, Patel, Mrunal, Chen, Han, Ferdman, Michael, Milder, Peter]
通讯作者:
Milder, Peter
SHF: Small: Massively Parallel Server Processors
-
批准号:2153297
-
项目类别:Standard Grant
-
资助金额:$59.88万
-
财政年份:2022
-
负责人:Michael Ferdman
-
依托单位:
FoMR: IPC Improvement through Hardware Memorization
-
批准号:1912517
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2019
-
负责人:Michael Ferdman
-
依托单位:
Student Travel - IEEE International Symposium on Workload Characterization (IISWC)
-
批准号:1737875
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2017
-
负责人:Michael Ferdman
-
依托单位:
SPX: Collaborative Research: Harnessing the Power of High-Bandwidth Memory via Provably Efficient Parallel Algorithms
-
批准号:1725543
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2017
-
负责人:Michael Ferdman
-
依托单位:
CAREER: Leveraging temporal streams for micro-architectural innovation in data center servers
-
批准号:1452904
-
项目类别:Continuing Grant
-
资助金额:$39.76万
-
财政年份:2015
-
负责人:Michael Ferdman
-
依托单位:
Preliminary Study to Demonstrate the Performance and Power Advantages of FPGAs over GPUs for Deep Learning in Computer Vision
-
批准号:1453460
-
项目类别:Standard Grant
-
资助金额:$9.5万
-
财政年份:2014
-
负责人:Michael Ferdman
-
依托单位:
II-New: Secure and Efficient Cloud Infrastructure and Accessibility Services
-
批准号:1405641
-
项目类别:Standard Grant
-
资助金额:$19.99万
-
财政年份:2014
-
负责人:Michael Ferdman
-
依托单位:
国内基金
海外基金
钴基Full-Heusler合金的掺杂效应和薄膜噪声特性研究
-
批准号:51871067
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2018
-
负责人:吴晟
-
依托单位: