High-performance, massively scalable distributed systems using the MapReduce software framework: the SHARD triple-store

High-performance, massively scalable distributed systems using the MapReduce software framework: the SHARD triple-store
复制标题

DOI:
10.1145/1940747.1940751
复制
发表时间:
2010-10
期刊:
--
影响因子:
--
通讯作者:
Kurt Rohloff;R. Schantz
Kurt Rohloff;R. Schantz
中科院分区:
其他
文献类型:
--
作者:
Kurt Rohloff;R. Schantz

文献摘要

被引文献

相似文献

在本文中,我们讨论了如何使用MapReduce软件框架来解决构建高性能、大规模可扩展的分布式系统的挑战。我们讨论了与使用MapReduceTM软件框架构建复杂的分布式系统相关的几个设计注意事项,包括可伸缩地构建索引的困难。我们关注的是Hadoop,这是最流行的MapReduce实现。我们的讨论和分析的动机是我们构建了Shard,这是一种基于Hadoop的大规模可扩展、高性能和健壮的三重存储技术。我们提供了一种从MapReduceTM软件框架构建响应数据查询的信息系统的通用方法。我们提供了一个早期版本的Shard生成的实验结果。最后,我们讨论了可用于构建更具可伸缩性的分布式计算系统的假想的MapReduceTM替代方案。
In this paper we discuss the use of the MapReduce software framework to address the challenge of constructing high-performance, massively-scalable distributed systems. We discuss several design considerations associated with constructing complex distributed systems using the MapReduce software framework, including the difficulty of scalably building indexes. We focus on Hadoop, the most popular MapReduce implementation. Our discussion and analysis are motivated by our construction of SHARD, a massively scalable, high-performance and robust triple-store technology on top of Hadoop. We provide a general approach to construct an information system from the MapReduce software framework that responds to data queries. We provide experimental results generated of an early version of SHARD. We close with a discussion of hypothetical MapReduce alternatives that can be used for the construction of more scalable distributed computing systems.