Fast crash recovery strategies for many small data objects in a distributed memory storageAkronym: FastRecovery
Fast crash recovery strategies for many small data objects in a distributed memory storageAkronym: FastRecovery
批准号:
269648469
负责人:
Professor Dr. Michael Schöttner
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2015
资助国家:
德国
项目状态:
已结题
起止时间:
2014-12-31 至 2017-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
More and more programs need to manage billions of small data objects like for example social network applications. Data access times of disks and solid state drives are too slow for these interactive applications and providers are forced to keep many data in caches. For large applications data cannot be loaded into memory of a single node and thus the memory of potentially many nodes need to be aggregated. A prominent example is Facebook running more than 1,000 memcached servers to keep around 75% of all data always in memory because background databases are too slow. Obviously, data is lost in case of node failures and power outages and it can take hours to load large data volumes from secondary storage like databases or file systems. The proposed project addresses these challenges by developing and evaluating fast recovery strategies for distributed memory systems. This project focuses on a key-value data-model for up to one trillion small data objects (sizes around 16-64 byte, stored in 1,000 nodes). Recovery will use an asynchronous logging strategy optimized for SSD drives, based on research from log-structured file systems. The state of one node needs to be distributed on many backup nodes in order to allow a fast and parallel recovery. All the log parts belonging to one node state need also to be replicated in order to be able to mask permanent node failures. It is important to point out that random replica placement has a high probability of data loss for large clusters, if several nodes fail simultaneously. We plan to address this challenge based upon the recently proposed Copyset replica placement scheme and we plan to develop efficient and adaptive strategies which minimize data loss probability while at the same time allow fast recovery. The backup management will be implemented using a super-peer overlay network taking into account different metrics including load and ongoing recoveries as well as re-replication.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
High Throughput Log-Based Replication for Many Small In-Memory Objects
针对许多小型内存对象的高吞吐量基于日志的复制
DOI:
10.1109/icpads.2016.0077
发表时间:
2016
期刊:
2016 IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
作者:
[Kevin Beineke, Stefan Nothaas, Michael Schöttner]
通讯作者:
Michael Schöttner
DOI:
10.1109/icpads.2017.00042
发表时间:
2017-12
期刊:
2017 IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
作者:
[Kevin Beineke;Stefan Nothaas;M. Schöttner]
通讯作者:
Kevin Beineke;Stefan Nothaas;M. Schöttner
Efficient Messaging for Java Applications Running in Data Centers
数据中心运行的 Java 应用程序的高效消息传递
DOI:
10.1109/ccgrid.2018.00090
发表时间:
2018
期刊:
2018 18th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)
影响因子:
--
作者:
[Kevin Beineke, Stefan Nothaas, Michael Schöttner]
通讯作者:
Michael Schöttner
海外基金