Exploring Transfer Learning to Reduce Training Overhead of HPC Data in Machine Learning

Exploring Transfer Learning to Reduce Training Overhead of HPC Data in Machine Learning
复制标题

DOI:
10.1109/nas.2019.8834723
复制
发表时间:
2019-08
期刊:
2019 IEEE International Conference on Networking, Architecture and Storage (NAS)
影响因子:
--
通讯作者:
Tong Liu;Shakeel Alibhai;Jinzhen Wang;Qing Liu;Xubin He;Chentao Wu
Tong Liu;Shakeel Alibhai;Jinzhen Wang;Qing Liu;Xubin He;Chentao Wu
中科院分区:
其他
文献类型:
--
作者:
Tong Liu;Shakeel Alibhai;Jinzhen Wang;Qing Liu;Xubin He;Chentao Wu

文献摘要

相似文献

如今,关于高性能计算(HPC)系统的科学模拟可以每次运行生成大量数据(以trabytes或pet的规模)。当通过机器学习应用程序处理大量的HPC数据时,培训开销将非常重要。通常,神经网络的培训过程可能需要几个小时才能完成,甚至不再长时间。当机器学习应用于HPC科学数据时,培训时间可能需要几天甚至几周。转移学习是一种通常用于节省训练时间或获得更好性能的优化,具有减少大型培训开销的潜力。在本文中,我们将转移学习应用于机器学习HPC应用程序。我们发现,转移学习可以减少训练时间,而在大多数情况下,可以大大增加错误。这表明转移学习对于在机器学习应用程序中使用HPC数据集非常有用。
Nowadays, scientific simulations on high-performance computing (HPC) systems can generate large amounts of data (in the scale of terabytes or petabytes) per run. When this huge amount of HPC data is processed by machine learning applications, the training overhead will be significant. Typically, the training process for a neural network can take several hours to complete, if not longer. When machine learning is applied to HPC scientific data, the training time can take several days or even weeks. Transfer learning, an optimization usually used to save training time or achieve better performance, has potential for reducing this large training overhead. In this paper, we apply transfer learning to a machine learning HPC application. We find that transfer learning can reduce training time without, in most cases, significantly increasing the error. This indicates transfer learning can be very useful for working with HPC datasets in machine learning applications.