2PCP: Two-phase CP decomposition for billion-scale dense tensors

2PCP: Two-phase CP decomposition for billion-scale dense tensors
复制标题

2PCP:十亿级稠密张量的两相CP分解

DOI:
10.1109/icde.2016.7498294
复制
发表时间:
2016
期刊:
2016 IEEE 32nd International Conference on Data Engineering (ICDE)
影响因子:
--
通讯作者:
M. Sapino
M. Sapino
中科院分区:
--
文献类型:
--
作者:
Xinsheng Li;Shengyu Huang;K. Candan;M. Sapino

文献摘要

被引文献

相似文献

张量是多维数组-因此,张量分解操作(CP和Tucker)是许多高维数据分析任务的基础,从聚类,趋势检测,异常检测到各种应用领域的相关性分析,包括科学和工程1。张量分解的一个关键问题是计算复杂度和空间要求。特别是,随着相关数据集越来越密集,张量分解的内存方案变得越来越无效;因此,需要核外(支持辅助内存,可能是并行的)计算。然而,现有的技术并没有考虑到在核外执行张量分解操作所带来的I/O和网络数据交换成本。在本文中,我们注意到,当在辅助内存和/或多个服务器的帮助下实现此操作以解决内存限制时,我们将需要智能缓冲区管理和任务调度技术,这些技术将考虑将相关块带入缓冲区的成本,以最小化系统中的I/O。本文介绍了一种基于智能缓冲区敏感任务调度和缓冲区管理机制的两阶段分块CP分解系统2PCP。2PCP旨在降低在科学和工程应用中常见的相对密集张量分析中的I/O成本。实验结果与目前最先进的张量分解算法进行了比较,表明我们的算法可以在保持分解精度的同时显着减少I/O量和执行时间。
Tensors are multi-dimensional arrays - consequently, tensor decomposition operations (CP and Tucker) are the bases for many high-dimensional data analysis tasks, from clustering, trend detection, anomaly detection, to correlation analysis in various application domains, including science and engineering1. One key problem with tensor decomposition is its computational complexity and space requirements. Especially, as the relevant data sets get denser, in-memory schemes for tensor decomposition become increasingly ineffective; therefore out-of-core (secondary-memory supported, potentially parallel) computing is necessitated. However, existing techniques do not consider the I/O and network data exchange costs that out-of-core execution of the tensor decomposition operation will incur. In this paper, we note that when this operation is implemented with the help of secondary-memory and/or multiple servers to tackle the memory limitations, we would need intelligent buffer-management and task-scheduling techniques which take into account the cost of bringing the relevant blocks into the buffer to minimize I/O in the system. In this paper, we introduce 2PCP, a two-phase, block-based CP decomposition system with intelligent buffer sensitive task scheduling and buffer management mechanisms. 2PCP aims to reduce I/O costs in the analysis of relatively dense tensors common in scientific and engineering applications. Experiment results compare with current state of art tensor decomposition algorithms and show that our algorithms can significantly reduce the amount of I/O and execution time while maintaining decomposition accuracy.