A Distributed Look-up Architecture for Text Mining Applications using MapReduce.

A Distributed Look-up Architecture for Text Mining Applications using MapReduce.
复制标题

DOI:
10.1145/1996130.1996174
复制
发表时间:
2011-11
期刊:
Proceedings of the ... International Symposium on High Performance Distributed Computing
影响因子:
--
通讯作者:
Rzhetsky A
Rzhetsky A
中科院分区:
其他
文献类型:
--
作者:
Balkir AS;Foster I;Rzhetsky A

文献摘要

相似文献

Text mining applications typically involve statistical models that require accessing and updating model parameters in an iterative fashion. With the growing size of the data, such models become extremely parameter rich, and naive parallel implementations fail to address the scalability problem of maintaining a distributed look-up table that maps model parameters to their values. We evaluate several existing alternatives to provide coordination among worker nodes in Hadoop clusters, and suggest a new multi-layered look-up architecture that is specifically optimized for certain problem domains. Our solution exploits the power-law distribution characteristics of the phrase or n-gram counts in large corpora while utilizing a Bloom Filter, in-memory cache, and an HBase cluster at varying levels of abstraction.