Data-Parallel Algorithms for Efficient Query Processing on Modern Hardware
Data-Parallel Algorithms for Efficient Query Processing on Modern Hardware
批准号:
RGPIN-2020-06639
负责人:
Chester, Sean
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
联合国强调了一个可持续发展目标(8.4),即到2030年提高全球资源利用效率,使经济增长与环境退化脱钩。数据分析是这种挑战的一个典型例子:它是知识经济的推动力,但它也在以指数级的规模增长。我们目前管理指数级增长的能源需求的解决方案是通过将分析转移到超大规模数据中心(如Amazon EC2)来提高效率。然而,根据国际能源署最新的数字化报告,到2020年,这一举措将完成一半(以太瓦时衡量);也就是说,这一能效提升的来源将很快耗尽。数据分析迫切需要另一个来源。今天的所有计算机都很复杂,包括那些形成超大规模云的计算机。然而,充分利用现代计算机是非常困难的:它自动重新排序计算,同时执行多条指令,并跨多个处理器和专门的加速器协调计算。此外,随着英特尔、AMD和英伟达等制造商不知疲倦地创新,“现代计算机”是一个不断变化的目标。因此,不足为奇的是,有一些计算任务--特别是那些涉及高价值文本、时空和社交数据源的任务--浪费了每台计算机内部用于并行计算的大部分机会。我们需要新的算法和数据结构,能够更有效地利用所有复杂的现代计算机,并可以在云中进行扩展,以便单个分析查询不会扩展到不必要的程度。也许更重要的是,我们需要具有足够先进技能的各种高素质人员,以将这些想法应用到更具资源效率和创新性的加拿大工业中。具体地说,该研究计划将设计新的算法和数据结构,以在现代计算机和专用图形处理单元(GPU)中普及对额外并行性的访问。我们将为文本创建并行友好的数据结构,以支持自然语言处理(NLP)的最新进展,以便并行分析和现代NLP可以共同发展。我们将为GPU上的图形处理定义一个更简单的计算范例,允许分析师在专注于更高级别的算法概念的同时受益于GPU的“延迟隐藏”。我们还将为跨服务器扩展的多维数据定义新的以GPU为中心的数据结构,以便传统数据分析可以更好地利用多个级别的并行性。总而言之,这项研究将拓宽哪些类型的数据可以有效地利用现代计算平台的范围。因此,科学家和行业专业人士都可以更快地产生更多的知识,拥有更多的数据,以更节约资源的方式。
英文摘要
The United Nations has highlighted a Sustainable Development Goal (8.4) to improve global resource efficiency and "decouple economic growth from environmental degradation" by 2030. Data analytics is a characteristic example of this challenge: it is an impetus of the knowledge economy, but it also is growing at an exponential scale. Our current solution to manage the energy demands of exponential growth is to gain efficiency by moving analytics into hyperscale data centres, like Amazon EC2. However, according to the International Energy Agency's latest digilisation report, this move will be half complete by 2020 (measured in terms of TWh); that is to say, this source for efficiency gains will soon be exhausted. Data analytics imminently needs another source. All computers today are complex, including those that form hyperscale clouds. Fully utilising a modern computer, however, is very difficult: it automatically reorders computations, executes multiple instructions simultaneously, and coordinates computation across multiple processors and specialised accelerators. Furthermore, a "modern computer" is a moving target, as manufacturers such as Intel, AMD, and Nvidia tirelessly innovate. Unsurprisingly, then, there are computational tasks---particularly those involving high-value text, spatio-temporal, and social data sources---that squander most of the opportunities inside each computer for parallel computing. We need novel algorithms and data structures that more efficiently utilise all of a complex, modern computer and can scale in the cloud so that individual analytics queries do not scale out so unnecessarily far. Perhaps even moreso, we need diverse highly qualified personnel with sufficiently advanced skills to apply these ideas to an even more resource-efficient and innovative Canadian industry. Concretely, this research program will design novel algorithms and data structures to democratise access to additional parallelism in modern computers and dedicated graphics processing units (GPUs). We will create parallel-friendly data structures for text that support recent advances in natural language processing (NLP) so that parallel analytics and modern NLP can co-develop. We will define a simpler computational paradigm for graph processing on GPUs that allows analysts to benefit from GPU "latency hiding" while focusing on higher-level algorithmic concepts. And we will define new GPU-centric data structures for multi-dimensional data that scales across servers so that traditional data analytics can better exploit multiple levels of parallelism. In all, this research will broaden the scope of what types of data can effectively leverage modern computing platforms. As a result, scientists and industry professionals alike can generate more knowledge faster, with more data, in a more resource-efficient manner.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data-Parallel Algorithms for Efficient Query Processing on Modern Hardware
-
批准号:RGPIN-2020-06639
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2022
-
负责人:Chester, Sean
-
依托单位:
Data-Parallel Algorithms for Efficient Query Processing on Modern Hardware
-
批准号:RGPIN-2020-06639
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2020
-
负责人:Chester, Sean
-
依托单位:
Data-Parallel Algorithms for Efficient Query Processing on Modern Hardware
-
批准号:DGECR-2020-00324
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2020
-
负责人:Chester, Sean
-
依托单位:
国内基金
海外基金
强流低能加速器束流损失机理的Parallel PIC/MCC算法与实现
-
批准号:11805229
-
项目类别:青年科学基金项目
-
资助金额:27.0万元
-
批准年份:2018
-
负责人:张青鵾
-
依托单位: