TOD: GPU-accelerated Outlier Detection via Tensor Operations

TOD: GPU-accelerated Outlier Detection via Tensor Operations
复制标题

DOI:
10.14778/3570690.3570703
复制
发表时间:
2021-10
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Yue Zhao;George H. Chen;Zhihao Jia
Yue Zhao;George H. Chen;Zhihao Jia
中科院分区:
其他
文献类型:
--
作者:
Yue Zhao;George H. Chen;Zhihao Jia

文献摘要

相似文献

离群值检测(OD)是一项关键的机器学习任务,用于发现罕见和异常的数据样本,具有许多时间关键型应用,如欺诈检测和入侵检测。在这项工作中,我们提出了TOD,这是第一个基于张量的系统,用于分布式多GPU机器上的高效和可扩展的离群值检测。TOD背后的一个关键思想是将复杂的OD应用程序分解为基本张量代数运算符的小集合。这种分解使TOD能够通过利用硬件和软件中深度学习基础设施的最新进展来加速OD计算。此外,为了在有限的设备内存的现代GPU上部署内存密集型OD应用程序,我们介绍了两个关键技术。首先,可证明的量化加速OD计算,并通过自动执行特定的浮点运算以较低的精度减少其内存占用,同时可证明地保证没有精度损失。其次,为了利用多个GPU的聚合计算资源和内存容量,我们引入了自动并行计算,它将OD计算分解为小批量,以便在单个GPU上顺序执行和跨多个GPU并行执行。TOD支持各种OD算法。对11个真实世界和3个合成OD数据集的评估表明,TOD平均比领先的基于CPU的OD系统PyOD快10.9倍(最大加速比为38.9倍),并且可以处理比现有基于GPU的OD系统更大的数据集。此外,TOD允许轻松集成新的OD运算符,从而实现新兴和尚未发现的OD算法的快速原型设计。
Outlier detection (OD) is a key machine learning task for finding rare and deviant data samples, with many time-critical applications such as fraud detection and intrusion detection. In this work, we propose TOD, the first tensor-based system for efficient and scalable outlier detection on distributed multi-GPU machines. A key idea behind TOD is decomposing complex OD applications into a small collection of basic tensor algebra operators. This decomposition enables TOD to accelerate OD computations by leveraging recent advances in deep learning infrastructure in both hardware and software. Moreover, to deploy memory-intensive OD applications on modern GPUs with limited on-device memory, we introduce two key techniques. First, provable quantization speeds up OD computations and reduces its memory footprint by automatically performing specific floating-point operations in lower precision while provably guaranteeing no accuracy loss. Second, to exploit the aggregated compute resources and memory capacity of multiple GPUs, we introduce automatic batching , which decomposes OD computations into small batches for both sequential execution on a single GPU and parallel execution across multiple GPUs. TOD supports a diverse set of OD algorithms. Evaluation on 11 real-world and 3 synthetic OD datasets shows that TOD is on average 10.9X faster than the leading CPU-based OD system PyOD (with a maximum speedup of 38.9X), and can handle much larger datasets than existing GPU-based OD systems. In addition, TOD allows easy integration of new OD operators, enabling fast prototyping of emerging and yet-to-discovered OD algorithms.