Neural Math Word Problem Solver with Reinforcement Learning

Neural Math Word Problem Solver with Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-08
期刊:
--
影响因子:
--
通讯作者:
Danqing Huang;Jing Liu;Chin-Yew Lin;Jian Yin
Danqing Huang;Jing Liu;Chin-Yew Lin;Jian Yin
中科院分区:
其他
文献类型:
--
作者:
Danqing Huang;Jing Liu;Chin-Yew Lin;Jian Yin

文献摘要

被引文献

相似文献

序列到序列模型已应用于求解数学问题。但是,我们的实验分析表明,该模型有两个缺点:(1)在错误的位置产生虚假数字本文,我们将副本和对准机制纳入序列模型(即CASS)来解决这些缺点,以训练我们的模型,我们将增强性学习直接优化解决方案的精度。最大似然估计的差异”,该问题在训练期间使用替代目标,而评估指标是解决方案准确性(非差异)时间。 (2)强化学习比最大的可能性更高的性能;最先进的结果。
Sequence-to-sequence model has been applied to solve math word problems. The model takes math problem descriptions as input and generates equations as output. The advantage of sequence-to-sequence model requires no feature engineering and can generate equations that do not exist in training data. However, our experimental analysis reveals that this model suffers from two shortcomings: (1) generate spurious numbers; (2) generate numbers at wrong positions. In this paper, we propose incorporating copy and alignment mechanism to the sequence-to-sequence model (namely CASS) to address these shortcomings. To train our model, we apply reinforcement learning to directly optimize the solution accuracy. It overcomes the “train-test discrepancy” issue of maximum likelihood estimation, which uses the surrogate objective of maximizing equation likelihood during training while the evaluation metric is solution accuracy (non-differentiable) at test time. Furthermore, to explore the effectiveness of our neural model, we use our model output as a feature and incorporate it into the feature-based model. Experimental results show that (1) The copy and alignment mechanism is effective to address the two issues; (2) Reinforcement learning leads to better performance than maximum likelihood on this task; (3) Our neural model is complementary to the feature-based model and their combination significantly outperforms the state-of-the-art results.