HPTA: High-performance text analytics
HPTA: High-performance text analytics
复制标题
DOI:
10.1109/bigdata.2016.7840632
复制
发表时间:
2016-12
期刊:
影响因子:
--
通讯作者:
Hans Vandierendonck;Karen L. Murphy;Mahwish Arif;Dimitrios S. Nikolopoulos
中科院分区:
文献类型:
--
作者:
Hans Vandierendonck;Karen L. Murphy;Mahwish Arif;Dimitrios S. Nikolopoulos
One of the main targets of data analytics is unstructured data, which primarily involves textual data. High-performance processing of textual data is non-trivial. We present the HPTA library for high-performance text analytics. The library helps programmers to map textual data to a dense numeric representation, which can be handled more efficiently. HPTA encapsulates three performance optimizations: (i) efficient memory management for textual data, (ii) parallel computation on associative data structures that map text to values and (iii) optimization of the type of associative data structure depending on the program context. We demonstrate that HPTA outperforms popular frameworks for text analytics such as scikit-learn.