An Architecture for Parallel Topic Models

An Architecture for Parallel Topic Models
复制标题

DOI:
10.14778/1920841.1920931
复制
发表时间:
2010-09-01
影响因子:
2.5
通讯作者:
Narayanamurthy, Shravan
Narayanamurthy, Shravan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Smola, Alexander;Narayanamurthy, Shravan

文献摘要

被引文献

相似文献

本文描述了一种用于工作站集群上潜在主题模型推理的高性能采样体系结构。我们的系统比以前的工作快了一个数量级,它能够处理数亿个文档和数千个主题。该算法依赖于一种新颖的通信结构,即使用分布式(键、值)存储来同步计算机之间的采样器状态。我们的体系结构完全避免了单独计算和同步阶段的需要。相反,同时使用磁盘、CPU和网络来实现高性能。我们表明,这种架构是完全通用的,它可以很容易地扩展到更复杂的潜在变量模型,如n-grams和层次结构。
This paper describes a high performance sampling architecture for inference of latent topic models on a cluster of workstations. Our system is faster than previous work by over an order of magnitude and it is capable of dealing with hundreds of millions of documents and thousands of topics.The algorithm relies on a novel communication structure, namely the use of a distributed (key, value) storage for synchronizing the sampler state between computers. Our architecture entirely obviates the need for separate computation and synchronization phases. Instead, disk, CPU, and network are used simultaneously to achieve high performance. We show that this architecture is entirely general and that it can be extended easily to more sophisticated latent variable models such as n-grams and hierarchies.