A Tensor-Based Approach for Big Data Representation and Dimensionality Reduction

A Tensor-Based Approach for Big Data Representation and Dimensionality Reduction
复制标题

基于张量的大数据表示和降维方法

DOI:
10.1109/tetc.2014.2330516
复制
发表时间:
2014-07-01
影响因子:
5.9
通讯作者:
Min, Geyong
Min, Geyong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kuang, Liwei;Hao, Fei;Min, Geyong

文献摘要

被引文献

相似文献

多样性和准确性是海量、异质数据的两个显著特点。如何用统一的方案高效地表示和处理大数据一直是一个巨大的挑战。本文提出了一个统一的张量模型来表示非结构化、半结构化和结构化数据。利用张量扩张算子,将各种类型的数据表示为子张量,然后合并为一个统一的张量。为了提取小而有价值的核张量,提出了一种增量高阶奇异值分解(IHOSVD)方法。通过递归地应用增量矩阵分解算法,IHOSVD能够更新正交基和计算新的核张量。本文从时间复杂度、内存使用量和近似精度三个方面对该方法进行了分析。实例研究表明,从包含18%元素的核集重建的近似数据总体上可以保证93%的准确率。理论分析和实验结果表明,所提出的统一张量模型和IHOSVD方法在大数据表示和降维方面是有效的。
Variety and veracity are two distinct characteristics of large-scale and heterogeneous data. It has been a great challenge to efficiently represent and process big data with a unified scheme. In this paper, a unified tensor model is proposed to represent the unstructured, semistructured, and structured data. With tensor extension operator, various types of data are represented as subtensors and then are merged to a unified tensor. In order to extract the core tensor which is small but contains valuable information, an incremental high order singular value decomposition (IHOSVD) method is presented. By recursively applying the incremental matrix decomposition algorithm, IHOSVD is able to update the orthogonal bases and compute the new core tensor. Analyzes in terms of time complexity, memory usage, and approximation accuracy of the proposed method are provided in this paper. A case study illustrates that approximate data reconstructed from the core set containing 18% elements can guarantee 93% accuracy in general. Theoretical analyzes and experimental results demonstrate that the proposed unified tensor model and IHOSVD method are efficient for big data representation and dimensionality reduction.