课题基金 / 基金详情

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures
III:中:异构架构上的大规模并行数据分析
批准号:
1763434
负责人:
Samuel Madden
金额:
$120.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2023-05-31

项目摘要

项目成果

Samuel Madden的其他基金

相似基金

相关文献

中文摘要
翻译
互联设备和互联网的兴起导致计算机必须存储和处理的数据量空前增长。数据量的巨大增长与对即时答案和交互式分析的需求不断增长相吻合。越来越多的公司需要销售、网络流量、异常和其他业务趋势的实时报告。互联网连接的设备,如汽车和工业工厂,需要更实时的数据分析。这些趋势意味着数据库和数据分析平台必须从机器上提供更快的性能,这一事实推动了对可扩展多节点处理系统的极大兴趣,如Hadoop和Spark,它们将大型数据集的处理分布在机器集群上。不幸的是,由于这些平台的设计方式,它们对每个节点上的硬件资源的利用率非常低,通常会产生比原始硬件低数千倍的单节点吞吐量。在这个项目中,将追求一个正交的方向;将建立一个名为Proteus的系统,该系统将获得尽可能充分利用硬件的性能,重点是产生一个充分利用所有可用计算资源的可扩展系统。如果成功,该项目将产生广泛的影响,因为数据库和数据密集型并行计算系统被世界各地数百万企业使用,无论是在现场还是在计算云中;更好地利用硬件的这些系统的优化实现将改善响应时间并降低硬件和能源成本,从而节省数十亿美元的成本。Proteus将在单个处理器上的多个核心上并行化,以及利用GPU和英特尔至强融核等众核系统的优势。此外,Proteus还将能够利用大型不同的硬件集群,但目标是在不放弃这种效率的情况下做到这一点,而不是接受分布式计算的低效率。为此,Proteus项目的研究将集中在四个关键领域:(1)开发单个数据库算法的优化实现,例如top-k排序,顺序扫描,随机查找,图形和GPU和CPU的机器学习算法。(2)构建成本模型,预测这些算法在异构体系结构上的性能。(3)开发抽象底层硬件细节的中间语言,隐藏这些不同平台的细微差别,但不放弃性能。(4)构建一个优化器,该优化器使用成本模型将这些计划放置到异构硬件组合中,以获得每个查询计划的最佳总体性能。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响评审标准进行评估,被认为值得支持。
英文摘要
The rise of connected devices and the Internet has led to unprecedented growth in the volumes of data that computers must store and process. This enormous growth in data volumes coincides with a growing demand for immediate answers and interactive analytics. Increasingly companies need real-time reports of sales, network traffic, anomalies, and other business trends. Internet-connected devices, like cars and industrial plants, demand even more real-time analysis of data. These trends mean database and data analytics platforms must deliver ever-faster performance from machines, a fact that has driven the dramatic interest in scalable multi-node processing systems, like Hadoop and Spark, which distribute the processing of large data sets across clusters of machines. Unfortunately, because of the way these platforms are engineered, they provide shockingly poor utilization of the hardware resources on each node, often times yielding single-node throughput that is thousands of times lower than what the raw hardware is capable of. In this project, an orthogonal direction will be pursued; a system, called Proteus, will be built that will obtain performance that utilizes hardware to the fullest extent possible, focusing on yielding a scalable system that fully utilizes all available computing resources. If successful, this project will have broad impact because databases and data-intensive parallel computing systems are used by millions of enterprises around the world, both on-site and in computing clouds; optimized implementations of these systems that better exploit hardware will improve response times and reduce hardware and energy costs, resulting in billions of dollars of cost savings.Proteus will parallelize across many cores on a single processor, as well as take advantages of many-core systems such as GPUs and Intel's Xeon Phi. In addition, Proteus will also be able to exploit large diverse clusters of hardware, but the aim is to do that without giving up this efficiency, rather than accepting inefficiency as a given of distributed computing. To do this, research in the Proteus project will focus on four key areas: (1) Developing optimized implementations of individual database algorithms, such as top-k sorts, sequential scans, random lookups, graph and machine learning algorithms for GPUs and CPUs. (2) Building cost models that predict the performance of these algorithms on heterogeneous architectures. (3) Developing intermediate languages that abstract details of the underlying hardware, to hide the nuances of these different platforms to but without giving up performance. (4) Building an optimizer that uses cost models to place these plans onto a heterogeneous mix of hardware to obtain the best overall performance for each query plan.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/isca52012.2021.00041
发表时间: 2021-06
期刊: 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者: [Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming]
通讯作者: Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming
DOI: 10.1145/3368826.3377909
发表时间: 2019-11
期刊: Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization
影响因子: --
作者: [Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun]
通讯作者: Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun
DOI: 10.1145/3297858.3304025
发表时间: 2019-04
期刊: Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者: [Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer]
通讯作者: Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer
DOI: 10.1145/3318464.3380595
发表时间: 2020-03
期刊: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子: --
作者: [Anil Shanbhag;S. Madden;Xiangyao Yu]
通讯作者: Anil Shanbhag;S. Madden;Xiangyao Yu
11
    Collaborative Research: Elements: A Self-tuning Anomaly Detection Service
    BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
    III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
    ACM SIGMOD 2012 Student Programming Contest: A Multidimensional Indexing System
    海外基金