HPTA: High-performance text analytics

HPTA: High-performance text analytics
复制标题

DOI:
10.1109/bigdata.2016.7840632
复制
发表时间:
2016-12
期刊:
2016 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Hans Vandierendonck;Karen L. Murphy;Mahwish Arif;Dimitrios S. Nikolopoulos
Hans Vandierendonck;Karen L. Murphy;Mahwish Arif;Dimitrios S. Nikolopoulos
中科院分区:
其他
文献类型:
--
作者:
Hans Vandierendonck;Karen L. Murphy;Mahwish Arif;Dimitrios S. Nikolopoulos

文献摘要

相似文献

数据分析的主要目标之一是非结构化数据,主要涉及文本数据。文本数据的高性能处理是非常重要的。我们提出了用于高性能文本分析的HPTA库。该库帮助程序员将文本数据映射为密集的数字表示,从而可以更有效地处理。HPTA封装了三个性能优化:(i)高效的文本数据的内存管理,(ii)关联数据结构的并行计算,将文本映射到值,以及(iii)根据程序上下文优化关联数据结构的类型。我们证明HPTA优于流行的文本分析框架,如scikit-learn。
One of the main targets of data analytics is unstructured data, which primarily involves textual data. High-performance processing of textual data is non-trivial. We present the HPTA library for high-performance text analytics. The library helps programmers to map textual data to a dense numeric representation, which can be handled more efficiently. HPTA encapsulates three performance optimizations: (i) efficient memory management for textual data, (ii) parallel computation on associative data structures that map text to values and (iii) optimization of the type of associative data structure depending on the program context. We demonstrate that HPTA outperforms popular frameworks for text analytics such as scikit-learn.