Combinatorial designs for deep learning

Combinatorial designs for deep learning
复制标题

DOI:
10.1002/jcd.21720
复制
发表时间:
2018-09
影响因子:
0.7
通讯作者:
S. Chisaki;R. Fuji-Hara;N. Miyamoto
S. Chisaki;R. Fuji-Hara;N. Miyamoto
中科院分区:
数学3区
文献类型:
--
作者:
S. Chisaki;R. Fuji-Hara;N. Miyamoto

文献摘要

相似文献

深度学习是一种使用多层神经网络的机器学习方法。设V1,V2,.,VL是相互不相交的节点集(层)。一个多层神经网络可以看作是两个连续节点集Vi和Vi+1上的完全二分图K <$Vi <$,<$Vi+1 <$的并集,其中i= 1,2,.,L −1。二分图的边作为权值,权值表示为矩阵。第i层的值基本上通过权重矩阵与第(i-1)层的值相乘来计算。使用大量的训练和教师数据,逐步估计权重参数。过度拟合(或过度学习)指的是一个模型,模型的“训练数据”太好。然后,模型很难推广到不在训练集中的新数据。避免过拟合的最流行的方法称为dropout。Dropout在训练过程中将激活(节点)的随机样本归零。节点的随机采样导致更不规则的丢弃边频率。在实验设计领域也有类似的抽样概念。我们提出了一个组合设计,从每一层的节点。这种设计平衡了边沿频率。我们在本文中分析和构造这样的设计。
Deep learning is a machine learning methodology using a multilayer neural network. Let V1,V2,…,VL be mutually disjoint node sets (layers). A multilayer neural network can be regarded as a union of the complete bipartite graphs K∣Vi∣,∣Vi+1∣ on consecutive two node sets Vi and Vi+1 for i=1,2,…,L−1 . The edges of a bipartite graph function as weights which are represented as a matrix. The values of i th layer are basically computed by multiplication of the weight matrix and values of (i−1) th layer. Using mass training and teacher data, the weight parameters are estimated little by little. Overfitting (or overlearning) refers to a model that models the “training data” too well. It then becomes difficult for the model to generalize to new data which were not in the training set. The most popular method to avoid overfitting is called dropout. Dropout zeros out a random sample of activations (nodes) during the training process. A random sampling of nodes causes more irregular frequency of dropout edges. There is a similar sampling concept in the area of design of experiments. We propose a combinatorial design that drops out nodes from each layer. This design balances the edge frequencies. We analyze and construct such designs in this paper.