Survey: Exploiting Data Redundancy for Optimization of Deep Learning

Survey: Exploiting Data Redundancy for Optimization of Deep Learning
复制标题

DOI:
10.1145/3564663
复制
发表时间:
2022-08
影响因子:
16.6
通讯作者:
Jou-An Chen;Wei Niu;Bin Ren;Yanzhi Wang;Xipeng Shen
Jou-An Chen;Wei Niu;Bin Ren;Yanzhi Wang;Xipeng Shen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jou-An Chen;Wei Niu;Bin Ren;Yanzhi Wang;Xipeng Shen

文献摘要

相似文献

数据冗余在深度神经网络(DNN)的输入和中间结果中无处不在。它为提高DNN的性能和效率提供了许多重要的机会,并且已经在大量的工作中进行了探索。这些研究分散在许多地点,跨越几年。他们关注的目标范围从图像到视频和文本,他们用来检测和利用数据冗余的技术也在许多方面有所不同。目前还没有一个系统的检查和总结的许多努力,使研究人员很难得到一个全面的看法,以前的工作,最先进的状态,差异和共同的原则,以及领域和方向尚未探索。本文试图填补这一空白。它调查了最近数百篇关于这个主题的论文,介绍了一种新的分类法,将各种技术放入一个单一的分类框架中,全面描述了用于利用数据冗余来改进多种DNN的主要方法,并指出了一系列未来探索的研究机会。
Data redundancy is ubiquitous in the inputs and intermediate results of Deep Neural Networks (DNN). It offers many significant opportunities for improving DNN performance and efficiency and has been explored in a large body of work. These studies have scattered in many venues across several years. The targets they focus on range from images to videos and texts, and the techniques they use to detect and exploit data redundancy also vary in many aspects. There is not yet a systematic examination and summary of the many efforts, making it difficult for researchers to get a comprehensive view of the prior work, the state of the art, differences and shared principles, and the areas and directions yet to explore. This article tries to fill the void. It surveys hundreds of recent papers on the topic, introduces a novel taxonomy to put the various techniques into a single categorization framework, offers a comprehensive description of the main methods used for exploiting data redundancy in improving multiple kinds of DNNs on data, and points out a set of research opportunities for future exploration.