Anomaly Detection in Scientific Datasets using Sparse Representation

Anomaly Detection in Scientific Datasets using Sparse Representation
复制标题

使用稀疏表示的科学数据集中的异常检测

DOI:
10.1145/3588982.3603610
复制
发表时间:
2023
期刊:
Proceedings of the First Workshop on AI for Systems
影响因子:
--
通讯作者:
Son, Seung Woo
Son, Seung Woo
中科院分区:
--
文献类型:
--
作者:
Moon, Aekyeung;Kim, Minjun;Chen, Jiaxi;Son, Seung Woo

文献摘要

参考文献

被引文献

相似文献

随着高性能计算 (HPC) 系统的规模和复杂性不断增长,科学家信任所生成数据的能力至关重要,因为各种原因可能导致数据损坏,而这些损坏可能未被发现。虽然采用基于机器学习的异常检测技术可以减轻科学家的这种担忧,但由于需要为大量科学数据集添加标签以及相关的不必要的额外开销,因此实际上是不可行的。在本文中,我们利用科学数据集中表现出的空间稀疏性分布,并提出了一种有效检测异常的方法。我们的方法首先在变换域中提取原始数据集的块级稀疏表示。然后,它从提取的稀疏表示中学习,并在不依赖标记数据的情况下构建正常和异常之间的边界阈值。使用真实世界科学数据集的实验表明,与两种最先进的无监督技术相比,所提出的方法平均需要整个数据集的 13%(大多数情况下小于 10%,低至 0.3%)才能实现有竞争力的检测精度 (70.74%-100.0%)。
As the size and complexity of high-performance computing (HPC) systems keep growing, scientists' ability to trust the data produced is paramount due to potential data corruption for various reasons, which may stay undetected. While employing machine learning-based anomaly detection techniques could relieve scientists of such concern, it is practically infeasible due to the need for labels for volumes of scientific datasets and the unwanted extra overhead associated. In this paper, we exploit spatial sparsity profiles exhibited in scientific datasets and propose an approach to detect anomalies effectively. Our method first extracts block-level sparse representations of original datasets in the transformed domain. Then it learns from the extracted sparse representations and builds the boundary threshold between normal and abnormal without relying on labeled data. Experiments using real-world scientific datasets show that the proposed approach requires 13% on average (less than 10% in most cases and as low as 0.3%) of the entire dataset to achieve competitive detection accuracy (70.74%-100.0%) as compared to two state-of-the-art unsupervised techniques.
DOI: --
发表时间: 2017
期刊: ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子: --
作者:
Panruo Wu;Nathan Debardeleben;Qiang Guan;S. Blanchard;Jieyang Chen;Dingwen Tao;Xin Liang;Kaiming Ouyang;Zizhong Chen
通讯作者: Zizhong Chen
DOI: --
发表时间: 2018
期刊: Sustainable Computing: Informatics and Systems
影响因子: --
作者:
Omer Subasi;S. Di;L. Bautista;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello
通讯作者: F. Cappello
了解有损压缩对物联网智能农场分析的影响
DOI: --
发表时间: 2017
期刊: 2017 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Aekyeung Moon;Jaeyoung Kim;Jialing Zhang;Hang Liu;S. Son
通讯作者: S. Son
基于神经网络的无声错误检测器
DOI: --
发表时间: 2018
期刊: IEEE International Conference on Cluster Computing
影响因子: --
作者:
Chen Wang;Nikoli Dryden;F. Cappello;M. Snir
通讯作者: M. Snir
空间支持向量回归检测百亿亿次时代的无声错误
DOI: --
发表时间: 2016
期刊: IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing
影响因子: --
作者:
Omer Subasi;S. Di;L. Bautista;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;F. Cappello
通讯作者: F. Cappello