课题基金 / 基金详情

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures

III: Medium: Massively Parallel Data Analytics on Heterogeneous Architectures
III:中:异构架构上的大规模并行数据分析
批准号:
1763434
负责人:
Samuel Madden
金额:
$120.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2023-05-31

项目摘要

项目成果

Samuel Madden的其他基金

相似基金

相关文献

中文摘要
翻译
连接设备和互联网的兴起导致计算机必须存储和处理的数据量空前增长。数据量的巨大增长与对即时答案和交互式分析的需求不断增长相一致。越来越多的公司需要销售、网络流量、异常情况和其他业务趋势的实时报告。汽车和工厂等联网设备需要更加实时的数据分析。这些趋势意味着数据库和数据分析平台必须为机器提供更快的性能,这一事实引发了人们对可扩展多节点处理系统的极大兴趣,例如 Hadoop 和 Spark,它们跨机器集群分布大型数据集的处理。不幸的是,由于这些平台的设计方式,它们对每个节点上的硬件资源的利用率极低,通常产生的单节点吞吐量比原始硬件的能力低数千倍。在这个项目中,将追求正交方向;将构建一个名为 Proteus 的系统,该系统将获得最大限度地利用硬件的性能,重点是产生一个可充分利用所有可用计算资源的可扩展系统。如果成功,该项目将产生广泛的影响,因为全球数百万企业在现场和计算云中使用数据库和数据密集型并行计算系统;这些系统的优化实现可以更好地利用硬件,从而缩短响应时间并降低硬件和能源成本,从而节省数十亿美元的成本。Proteus 将在单个处理器上的多个内核上进行并行化,并利用 GPU 和英特尔至强融核等多核系统的优势。此外,Proteus 还能够利用大型多样化的硬件集群,但目标是在不放弃这种效率的情况下做到这一点,而不是接受分布式计算的低效率。为此,Proteus 项目的研究将集中在四个关键领域:(1)开发单个数据库算法的优化实现,例如用于 GPU 和 CPU 的 top-k 排序、顺序扫描、随机查找、图形和机器学习算法。 (2) 构建成本模型来预测这些算法在异构架构上的性能。 (3) 开发抽象底层硬件细节的中间语言,隐藏这些不同平台的细微差别,但不牺牲性能。 (4) 构建一个优化器,使用成本模型将这些计划放置在异构硬件组合上,以获得每个查询计划的最佳整体性能。该奖项反映了 NSF 的法定使命,并通过使用基金会的智力优点和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The rise of connected devices and the Internet has led to unprecedented growth in the volumes of data that computers must store and process. This enormous growth in data volumes coincides with a growing demand for immediate answers and interactive analytics. Increasingly companies need real-time reports of sales, network traffic, anomalies, and other business trends. Internet-connected devices, like cars and industrial plants, demand even more real-time analysis of data. These trends mean database and data analytics platforms must deliver ever-faster performance from machines, a fact that has driven the dramatic interest in scalable multi-node processing systems, like Hadoop and Spark, which distribute the processing of large data sets across clusters of machines. Unfortunately, because of the way these platforms are engineered, they provide shockingly poor utilization of the hardware resources on each node, often times yielding single-node throughput that is thousands of times lower than what the raw hardware is capable of. In this project, an orthogonal direction will be pursued; a system, called Proteus, will be built that will obtain performance that utilizes hardware to the fullest extent possible, focusing on yielding a scalable system that fully utilizes all available computing resources. If successful, this project will have broad impact because databases and data-intensive parallel computing systems are used by millions of enterprises around the world, both on-site and in computing clouds; optimized implementations of these systems that better exploit hardware will improve response times and reduce hardware and energy costs, resulting in billions of dollars of cost savings.Proteus will parallelize across many cores on a single processor, as well as take advantages of many-core systems such as GPUs and Intel's Xeon Phi. In addition, Proteus will also be able to exploit large diverse clusters of hardware, but the aim is to do that without giving up this efficiency, rather than accepting inefficiency as a given of distributed computing. To do this, research in the Proteus project will focus on four key areas: (1) Developing optimized implementations of individual database algorithms, such as top-k sorts, sequential scans, random lookups, graph and machine learning algorithms for GPUs and CPUs. (2) Building cost models that predict the performance of these algorithms on heterogeneous architectures. (3) Developing intermediate languages that abstract details of the underlying hardware, to hide the nuances of these different platforms to but without giving up performance. (4) Building an optimizer that uses cost models to place these plans onto a heterogeneous mix of hardware to obtain the best overall performance for each query plan.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/isca52012.2021.00041
发表时间: 2021-06
期刊: 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者: [Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming]
通讯作者: Ajay Brahmakshatriya;Emily Furst;Victor A. Ying;Claire Hsu;Changwan Hong;Max Ruttenberg;Yunming Zhang-Yunming
DOI: 10.1145/3368826.3377909
发表时间: 2019-11
期刊: Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization
影响因子: --
作者: [Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun]
通讯作者: Yunming Zhang;Ajay Brahmakshatriya;Xinyi Chen;Laxman Dhulipala;Shoaib Kamil;Saman P. Amarasinghe;Julian Shun
DOI: 10.1145/3297858.3304025
发表时间: 2019-04
期刊: Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者: [Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer]
通讯作者: Michael Pellauer;Y. Shao;Jason Clemons;N. Crago;Kartik Hegde;Rangharajan Venkatesan;S. Keckler;Christopher W. Fletcher;J. Emer
DOI: 10.1145/3318464.3380595
发表时间: 2020-03
期刊: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子: --
作者: [Anil Shanbhag;S. Madden;Xiangyao Yu]
通讯作者: Anil Shanbhag;S. Madden;Xiangyao Yu
11
    Collaborative Research: Elements: A Self-tuning Anomaly Detection Service
    BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
    III: Medium: Collaborative Research: DataHub - A Collaborative Dataset Management Platform for Data Science
    ACM SIGMOD 2012 Student Programming Contest: A Multidimensional Indexing System
    海外基金