课题基金 / 基金详情

Scalable and Consistent Management of Large Scale Data

Scalable and Consistent Management of Large Scale Data
大规模数据的可扩展且一致的管理
批准号:
RGPIN-2014-03670
负责人:
Daudjee, Khuzaima
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31

项目摘要

项目成果

Daudjee, Khuzaima的其他基金

相似基金

相关文献

中文摘要
翻译
有效的大规模数据管理对于科学发现以及保持商业和信息技术的竞争优势变得越来越重要。大规模数据管理,有时称为“大数据”管理,通常用于指数据量及其生成速度。 CERN 等科学研究组织以及 Twitter、沃尔玛和 Facebook 等公司都生成和管理大量数据。例如,从 2016 年开始,大型综合巡天望远镜将在 10 年内每晚收集 30 TB 的数据。 Facebook 每天上传超过 100 TB 的数据,该公司在全球拥有超过 10 亿活跃用户。为了管理这些大规模数据,需要在地理上分布的数据中心内部和跨多个大型存储和处理服务器云进行大规模分布。此外,需要对这些服务器云进行有效管理,以降低运营成本。 为此,需要解决大规模数据管理的关键研究问题。首先,服务器云内部和跨服务器云都需要动态数据库分发技术来处理导致响应时间的延迟。这些延迟可能是由于工作负载请求的整体变化模式或由于负载峰值而出现的数据库热点造成的。其次,随着数据中心的增长,需要节省存储和管理大规模数据的服务器维护成本。因此,整合数据库服务器负载的技术成为降低成本和保持大规模数据管理可行的关键。 该提案的目标是设计和开发一套协议、算法和系统,利用动态、即时、大规模数据管理技术为这些研究问题提供可行的解决方案。这些技术将推进分布式数据管理的最先进水平,并为系统设计者和管理员提供“旋钮”,可以控制这些“旋钮”来改变降低管理数据成本所需的动态程度,同时量化性能权衡。
英文摘要
Effective large scale data management is becoming important for scientific discovery and for maintaining a competitive edge in business and information technology. Large scale data management, sometimes called “Big Data” management, is often used to refer to the volume of data and the rate at which it is generated. Scientific research organizations like CERN and companies such as Twitter, Walmart and Facebook all generate and manage very large amounts of data. For example, starting 2016, the Large Synoptic Survey Telescope will collect 30 terabytes of data every night for 10 years. Over 100 terabytes of data are uploaded daily to Facebook, which has more than 1 billion active users around the globe. To manage these large scale data, massive distribution over multiple, large, storage and processing server clouds both within and across geographically distributed data centers is required. Furthermore, effective management of these server clouds is needed so that operational costs are reduced. To this end, key research problems need to be addressed for large scale data management. First, dynamic database distribution techniques are required both within and across server clouds to deal with latencies contributing to response times. These latencies can be due to, for example, overall changing patterns in workload requests or the emergence of database hot spots as a result of load spikes. Second, with the growth of data centers, there is a need to achieve savings in terms of server maintenance cost of storing and managing large scale data. Thus, techniques to consolidate database server loads become key to reduce costs and keep large scale data management viable. The goal of this proposal is to design and develop a set of protocols, algorithms and systems that will deliver viable solutions to these research problems using dynamic, on-the-fly, large scale data management techniques. These techniques will advance the state-of-the-art in distributed data management and provide systems designers and administrators with "knobs" that can be controlled to vary the degree of dynamism required to reduce the costs of managing the data while quantifying performance trade-offs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adaptive Data Systems
  • 批准号:
    RGPIN-2019-05630
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2022
  • 负责人:
    Daudjee, Khuzaima
  • 依托单位:
Dynamic Partitioning for Partially Replicated Databases
  • 批准号:
    543858-2019
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $5.77万
  • 财政年份:
    2021
  • 负责人:
    Daudjee, Khuzaima
  • 依托单位:
Adaptive Data Systems
  • 批准号:
    RGPIN-2019-05630
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2021
  • 负责人:
    Daudjee, Khuzaima
  • 依托单位:
Dynamic Partitioning for Partially Replicated Databases
  • 批准号:
    543858-2019
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $5.77万
  • 财政年份:
    2020
  • 负责人:
    Daudjee, Khuzaima
  • 依托单位:
海外基金