An overview of gradient descent optimization algorithms

An overview of gradient descent optimization algorithms
复制标题

DOI:
10.14489/vkit.2019.12.pp.010-017
复制
发表时间:
2016-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Sebastian Ruder
Sebastian Ruder
中科院分区:
其他
文献类型:
--
作者:
Sebastian Ruder

文献摘要

被引文献

相似文献

统计测试技术被认为是在测试集上比较机器学习模型的度量值。由于度量的值不仅取决于模型,还取决于数据,因此可能会发现不同的模型在不同的测试集上是最好的。出于这个原因,比较测试集上的度量值的传统方法通常是不够的。有时使用基于交叉验证获得的结果的统计比较,但在这种情况下,不可能保证获得的测量的独立性,这不允许使用学生t检验。有一些标准不需要独立的测量,但它们的功率较小。对于可加性度量,本文提出了一种技术,当测试样本被分成N个部分,在每个部分上计算度量值。由于每个部分上的值是作为独立随机变量的和获得的,根据中心极限定理,N个部分中的每个部分上获得的度量值是正态分布随机变量的实现。为了估计所需的样本量,建议使用正态性检验并构建分位数-分位数图。然后,您可以使用Student t检验的修改来执行统计检验,比较度量的平均值。还考虑了一种简化的方法,其中置信区间为基础模型。度量值不在此区间内的模型的工作方式与基本模型不同。这种方法减少了所需的计算量,然而,CTR(点击率)预测模型的二进制交叉熵度量的实验分析表明,它比第一个更粗糙。
The statistical testing technique is considered to compare the metrics values of machine learning models on a test set. Since the values of metrics depend not only on the models, but also on the data, it may turn out that different models are the best on different test sets. For this reason, the traditional approach to comparing the values of metrics on a test set is often not enough. Sometimes a statistical comparison of the results obtained on the basis of cross-validation is used, but in this case it is impossible to guarantee the independence of the obtained measurements, which does not allow the use of the Student's t-test. There are criteria that do not require independent measurements, but they have less power. For additive metrics, a technique is proposed in this paper, when a test sample is divided into N parts, on each of which the values of the metrics are calculated. Since the value on each part is obtained as the sum of independent random variables, according to the central limit theorem, the obtained metrics values on each of the N parts are realizations of the normally distributed random variable. To estimate the required sample size, it is proposed to use normality tests and build quantile– quantile plots. You can then use a modification of the Student's t-test to conduct a statistical test comparing the mean values of the metrics. A simplified approach is also considered, in which confidence intervals are built for the base model. A model whose metric values do not fall into this interval works differently from the base model. This approach reduces the amount of computations needed, however, an experimental analysis of the binary cross-entropy metric for CTR (Click-Through Rate) prediction models showed that it is more rough than the first one.