Neural networks designing neural networks: Multi-objective hyper-parameter optimization

Neural networks designing neural networks: Multi-objective hyper-parameter optimization
复制标题

神经网络设计神经网络:多目标超参数优化

DOI:
--
复制
发表时间:
2016
期刊:
2016 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
B. Meyer
B. Meyer
中科院分区:
--
文献类型:
--
作者:
S. C. Smithson;Guang Yang;W. Gross;B. Meyer

文献摘要

被引文献

相似文献

人工神经网络最近越来越受欢迎,在各个领域取得了最先进的成果,包括图像分类,语音识别和自动控制。这种模型的性能和计算复杂度都严重依赖于特征超参数的设计(例如,隐藏层的数量、每层的节点、或者激活函数的选择),这些传统上是手动优化的。随着机器学习渗透到低功耗移动的和嵌入式领域,不仅需要优化性能(准确性),还需要优化实现复杂性,这变得至关重要。在这项工作中,我们提出了一种多目标设计空间探索方法,减少了通过响应面建模训练和评估的解决方案网络的数量。鉴于空间很容易超过1020个解决方案,手动设计接近最佳的架构是不太可能的,因为在保持性能的同时降低网络复杂性的机会可能会被忽视。在特定数据集上表现良好的超参数可能会在其他数据集上产生低于标准的结果,因此必须在每个应用程序的基础上设计,这一事实加剧了这个问题。在我们的工作中,机器学习通过训练人工神经网络来预测未来候选网络的性能。该方法在MNIST和CIFAR-10图像数据集上进行了评估,优化了识别精度和计算复杂度。实验结果表明,该方法可以接近帕累托最优前沿,而只有探索一小部分的设计空间。
Artificial neural networks have gone through a recent rise in popularity, achieving state-of-the-art results in various fields, including image classification, speech recognition, and automated control. Both the performance and computational complexity of such models are heavily dependant on the design of characteristic hyper-parameters (e.g., number of hidden layers, nodes per layer, or choice of activation functions), which have traditionally been optimized manually. With machine learning penetrating low-power mobile and embedded areas, the need to optimize not only for performance (accuracy), but also for implementation complexity, becomes paramount. In this work, we present a multi-objective design space exploration method that reduces the number of solution networks trained and evaluated through response surface modelling. Given spaces which can easily exceed 1020 solutions, manually designing a near-optimal architecture is unlikely as opportunities to reduce network complexity, while maintaining performance, may be overlooked. This problem is exacerbated by the fact that hyper-parameters which perform well on specific datasets may yield sub-par results on others, and must therefore be designed on a per-application basis. In our work, machine learning is leveraged by training an artificial neural network to predict the performance of future candidate networks. The method is evaluated on the MNIST and CIFAR-10 image datasets, optimizing for both recognition accuracy and computational complexity. Experimental results demonstrate that the proposed method can closely approximate the Pareto-optimal front, while only exploring a small fraction of the design space.