Why ADAGRAD Fails for Online Topic Modeling

Why ADAGRAD Fails for Online Topic Modeling
复制标题

DOI:
10.18653/v1/d17-1046
复制
发表时间:
2017-09
期刊:
--
影响因子:
--
通讯作者:
You Lu;Jeffrey Lund;Jordan L. Boyd-Graber
You Lu;Jeffrey Lund;Jordan L. Boyd-Graber
中科院分区:
其他
文献类型:
--
作者:
You Lu;Jeffrey Lund;Jordan L. Boyd-Graber

文献摘要

被引文献

相似文献

在线主题建模,即,具有随机变分推理的主题建模是用于分析大型数据集的强大且有效的技术,并且ADAGRAD是用于在线梯度优化期间调整学习速率的广泛使用的技术。然而,这两种技术并不能很好地结合在一起。我们表明,这是因为ADAGRAD使用先前梯度的累积作为学习率的修正器。对于在线主题建模,梯度的大小非常大。它会导致学习率非常快地收缩,因此在训练结束之前,参数无法完全收敛
Online topic modeling, i.e., topic modeling with stochastic variational inference, is a powerful and efficient technique for analyzing large datasets, and ADAGRAD is a widely-used technique for tuning learning rates during online gradient optimization. However, these two techniques do not work well together. We show that this is because ADAGRAD uses accumulation of previous gradients as the learning rates’ denominators. For online topic modeling, the magnitude of gradients is very large. It causes learning rates to shrink very quickly, so the parameters cannot fully converge until the training ends