Ringo: Interactive Graph Analytics on Big-Memory Machines.

Ringo: Interactive Graph Analytics on Big-Memory Machines.
复制标题

DOI:
10.1145/2723372.2735369
复制
发表时间:
2015-05
期刊:
Proceedings. ACM-SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Leskovec J
Leskovec J
中科院分区:
其他
文献类型:
--
作者:
Perez Y;Sosič R;Banerjee A;Puttagunta R;Raison M;Shah P;Leskovec J

文献摘要

被引文献

相似文献

我们提出了Ringo,一个大型图的分析系统。图形提供了一种表示和分析交互对象(人,蛋白质,网页)系统的方法,对象之间的边表示交互(友谊,物理交互,链接)。挖掘图提供了关于单个对象以及它们之间关系的有价值的见解。在构建Ringo时,我们利用了这样一个事实,即具有大内存和许多内核的机器广泛可用,并且相对便宜。这使我们能够构建一个易于使用的交互式高性能图形分析系统。图还需要从输入数据构建,输入数据通常以关系表的形式存在。因此,Ringo提供了丰富的功能,用于将原始输入数据表操作为各种图形。此外,Ringo还提供了200多个图形分析功能,可以应用于构建的图形。我们表明,一个大内存的机器提供了一个非常有吸引力的平台,除了最大的图形上执行分析,因为它提供了出色的性能和易用性相比,替代方法。通过Ringo,我们还演示了如何将图分析与数据挖掘工作负载中常见的反复试验数据探索和快速实验过程相集成。
We present Ringo, a system for analysis of large graphs. Graphs provide a way to represent and analyze systems of interacting objects (people, proteins, webpages) with edges between the objects denoting interactions (friendships, physical interactions, links). Mining graphs provides valuable insights about individual objects as well as the relationships among them. In building Ringo, we take advantage of the fact that machines with large memory and many cores are widely available and also relatively affordable. This allows us to build an easy-to-use interactive high-performance graph analytics system. Graphs also need to be built from input data, which often resides in the form of relational tables. Thus, Ringo provides rich functionality for manipulating raw input data tables into various kinds of graphs. Furthermore, Ringo also provides over 200 graph analytics functions that can then be applied to constructed graphs. We show that a single big-memory machine provides a very attractive platform for performing analytics on all but the largest graphs as it offers excellent performance and ease of use as compared to alternative approaches. With Ringo, we also demonstrate how to integrate graph analytics with an iterative process of trial-and-error data exploration and rapid experimentation, common in data mining workloads.