课题基金 / 基金详情

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures
III:中:异构架构上的大规模并行数据分析
批准号:
1763434
负责人:
Samuel Madden
金额:
$120.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2023-05-31

项目摘要

项目成果

Samuel Madden的其他基金

相似基金

相关文献

中文摘要
翻译
互联设备和互联网的兴起导致计算机必须存储和处理的数据量出现了前所未有的增长。数据量的巨大增长与对即时答案和交互分析的需求不断增长不谋而合。越来越多的公司需要实时报告销售、网络流量、异常情况和其他业务趋势。汽车和工业厂房等联网设备甚至需要对数据进行更实时的分析。这些趋势意味着数据库和数据分析平台必须提供更快的机器性能,这一事实促使人们对Hadoop和Spark等可扩展的多节点处理系统产生了极大的兴趣,这些系统将大型数据集的处理分布在多个机器群集之间。不幸的是,由于这些平台的设计方式,它们对每个节点上的硬件资源的利用率低得令人震惊,经常会产生比原始硬件所能达到的吞吐量低数千倍的单节点吞吐量。在这个项目中,将追求一个正交方向;将建造一个名为Proteus的系统,该系统将获得最大限度地利用硬件的性能,重点是产生一个充分利用所有可用计算资源的可扩展系统。如果成功,这个项目将产生广泛的影响,因为全球数以百万计的企业在现场和计算云中使用数据库和数据密集型并行计算系统;对这些系统进行优化实施,更好地利用硬件,将改善响应时间,降低硬件和能源成本,从而节省数十亿美元的成本。Proteus将在单个处理器上的多个内核上并行,并利用多核系统,如GPU和英特尔的Xeon Phi。此外,Proteus还将能够利用大型不同的硬件集群,但目标是在不放弃这种效率的情况下做到这一点,而不是接受低效率作为分布式计算的既定条件。为此,Proteus项目的研究将集中在四个关键领域:(1)开发单个数据库算法的优化实现,例如针对GPU和CPU的top-k排序、顺序扫描、随机查找、图形和机器学习算法。(2)建立代价模型来预测这些算法在异构性体系结构上的性能。(3)开发抽象底层硬件细节的中间语言,以隐藏这些不同平台的细微差别,但不会放弃性能。(4)构建一个优化器,它使用成本模型将这些计划放在不同的硬件组合上,以获得每个查询计划的最佳整体性能。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The rise of connected devices and the Internet has led to unprecedented growth in the volumes of data that computers must store and process. This enormous growth in data volumes coincides with a growing demand for immediate answers and interactive analytics. Increasingly companies need real-time reports of sales, network traffic, anomalies, and other business trends. Internet-connected devices, like cars and industrial plants, demand even more real-time analysis of data. These trends mean database and data analytics platforms must deliver ever-faster performance from machines, a fact that has driven the dramatic interest in scalable multi-node processing systems, like Hadoop and Spark, which distribute the processing of large data sets across clusters of machines. Unfortunately, because of the way these platforms are engineered, they provide shockingly poor utilization of the hardware resources on each node, often times yielding single-node throughput that is thousands of times lower than what the raw hardware is capable of. In this project, an orthogonal direction will be pursued; a system, called Proteus, will be built that will obtain performance that utilizes hardware to the fullest extent possible, focusing on yielding a scalable system that fully utilizes all available computing resources. If successful, this project will have broad impact because databases and data-intensive parallel computing systems are used by millions of enterprises around the world, both on-site and in computing clouds; optimized implementations of these systems that better exploit hardware will improve response times and reduce hardware and energy costs, resulting in billions of dollars of cost savings.Proteus will parallelize across many cores on a single processor, as well as take advantages of many-core systems such as GPUs and Intel's Xeon Phi. In addition, Proteus will also be able to exploit large diverse clusters of hardware, but the aim is to do that without giving up this efficiency, rather than accepting inefficiency as a given of distributed computing. To do this, research in the Proteus project will focus on four key areas: (1) Developing optimized implementations of individual database algorithms, such as top-k sorts, sequential scans, random lookups, graph and machine learning algorithms for GPUs and CPUs. (2) Building cost models that predict the performance of these algorithms on heterogeneous architectures. (3) Developing intermediate languages that abstract details of the underlying hardware, to hide the nuances of these different platforms to but without giving up performance. (4) Building an optimizer that uses cost models to place these plans onto a heterogeneous mix of hardware to obtain the best overall performance for each query plan.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/isca52012.2021.00041
发表时间: 2021-06
期刊: 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者: [Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming]
通讯作者: Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming
DOI: 10.1145/3368826.3377909
发表时间: 2019-11
期刊: Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization
影响因子: --
作者: [Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun]
通讯作者: Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun
DOI: 10.1145/3297858.3304025
发表时间: 2019-04
期刊: Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者: [Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer]
通讯作者: Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer
DOI: 10.1145/3318464.3380595
发表时间: 2020-03
期刊: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子: --
作者: [Anil Shanbhag;S. Madden;Xiangyao Yu]
通讯作者: Anil Shanbhag;S. Madden;Xiangyao Yu
11
    Collaborative Research: Elements: A Self-tuning Anomaly Detection Service
    BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
    III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
    ACM SIGMOD 2012 Student Programming Contest: A Multidimensional Indexing System
    海外基金