课题基金 / 基金详情

Efficient query processing and optimizations for big data workloads

Efficient query processing and optimizations for big data workloads
针对大数据工作负载的高效查询处理和优化
批准号:
RGPIN-2015-04587
负责人:
Koudas, Nikolaos
金额:
$4.37万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31

项目摘要

项目成果

Koudas, Nikolaos的其他基金

相似基金

相关文献

中文摘要
翻译
从感官数据采集吞吐量到处理器功率、存储和带宽,计算的各个方面都在经历指数级增长。这些指数级的改进正在推动大数据革命。大数据应用包括以流方式不断产生的大量数据(例如,传感器读数、日志、点击等)。此外,典型的大数据研究和分析工作流程是迭代的。也就是说,使用一些数据参数构建模型,然后使用前一个建模阶段的输出迭代地改进模型。这两种原语,即流数据生成和迭代分析工作流,都为优化提供了很多机会。******该项目的目标是探索这些原语,并提供基本的算法和技术,以有效地处理和优化大数据工作负载。最终目标是将这些技术包含到端到端数据处理体系结构中。流数据生成提供了以增量方式维护已经在数据上计算的模型的机会。作为第一个示例,可以为数据集中附加的新数据增量地维护在数据集中计算的统计运算符。此外,可以将已经在数据上计算的模型(它们之间或与基础数据)增量地组合起来,以计算新的建模查询请求的答案。就性能而言,结合两个模型可能比从头开始计算一个新模型要优越得多。******在这个项目中,我们计划在我们的系统设计中引入增量计算作为头等公民。我们将逐步维护感兴趣的模型(当新数据到达时);通过这些模型的具体化,分析阶段将能够重用从先前的分析中获得的结果。将开发适当的优化框架,以评估何时和在何种条件下这种组合和模型的增量维护是有益的。很明显,后续分析任务的性能将受益于模型重用和/或广泛的模型类别的组合,这些模型探索精确和近似计算。第二,我们计划建立一个包含我们创新的端到端系统。我们的设计将集中在统计处理和数据分析的流行语言上,以表达建模工作负载(例如,R)和合适的系统基础设施来实现和执行我们的框架。******我们研究的最终产品将是一个包含所有研究的系统,利用熟悉的分析查询处理接口(如r)提供非常快速的大数据分析,这样的系统将有利于并帮助数据科学家在所需的一小部分时间内进行高级研究,通过能够以增量方式无缝重用和共享结果。**
英文摘要
Every aspect of computing has been experiencing exponential growth, from sensory data acquisition throughput to processor power, storage and bandwidth. These exponential improvements are enabling the big data revolution. Big data applications consist of volumes of data that are constantly produced in a streaming fashion (e.g., sensor readings, logs, click-through etc.). In addition typical research and analysis workflows on big data are iterative. Namely a model is built using some data parameters, then iteratively refined using the output of the previous modeling phase. Both such primitives, namely streaming data generation and iterative analysis workflows, provide a lot of opportunity for optimizations. ******The goal of this project is to explore these primitives and deliver fundamental algorithms and techniques to efficiently process and optimize big data workloads. The end goal is to encompass such techniques into end-to-end data processing architecture. Streaming data generation provides the opportunity to maintain models already computed on the data in an incremental fashion. As a first example, a statistical operator computed on a data set can be incrementally maintained for new data appended in the data set. In addition, models already computed on the data can be combined (among themselves or with base data) incrementally to compute answers to new modeling query requests. Combining two models could be vastly superior in terms of performance than computing a new model from scratch. ******In this project we plan to introduce incremental computations as a first class citizen in our system design. We will incrementally maintain models of interest (as new data arrive); via materialization of such models, analysis phases will be able to re-use results available from prior analysis. Suitable optimization frameworks will be developed to assess when and under what conditions such combinations and incremental maintenance of models is beneficial. It is evident that the performance of subsequent analysis tasks will benefit from model re-use and/or combination for a wide class of models exploring both exact and approximate computations. Second, we plan to build an end-to-end system encompassing our innovations. Our design will be centered on popular languages for statistical processing and data analysis to express modeling workloads (e.g., R) and the suitable systems infrastructure to implement and execute our framework. ******The end product of our research will be a system encompassing all of the research conducted delivering very fast big data analytics utilizing familiar analytical query processing interfaces such as R. Such a system will benefit and help data scientists conduct advanced research in a fraction of the time required, by being able to seamlessly re-use and share results in an incremental fashion.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Declarative Query Processing Over Real Time Video Streams
  • 批准号:
    RGPIN-2020-07238
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2022
  • 负责人:
    Koudas, Nikolaos
  • 依托单位:
Declarative Query Processing Over Real Time Video Streams
  • 批准号:
    RGPIN-2020-07238
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2021
  • 负责人:
    Koudas, Nikolaos
  • 依托单位:
Declarative Query Processing Over Real Time Video Streams
  • 批准号:
    RGPIN-2020-07238
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.11万
  • 财政年份:
    2020
  • 负责人:
    Koudas, Nikolaos
  • 依托单位:
Efficient query processing and optimizations for big data workloads
  • 批准号:
    RGPIN-2015-04587
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $4.37万
  • 财政年份:
    2018
  • 负责人:
    Koudas, Nikolaos
  • 依托单位:
海外基金