Measuring Mathematical Problem Solving With the MATH Dataset

Measuring Mathematical Problem Solving With the MATH Dataset
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Dan Hendrycks;Collin Burns;Saurav Kadavath;Akul Arora;Steven Basart;Eric Tang;D. Song;J. Steinhardt
Dan Hendrycks;Collin Burns;Saurav Kadavath;Akul Arora;Steven Basart;Eric Tang;D. Song;J. Steinhardt
中科院分区:
其他
文献类型:
--
作者:
Dan Hendrycks;Collin Burns;Saurav Kadavath;Akul Arora;Steven Basart;Eric Tang;D. Song;J. Steinhardt

文献摘要

被引文献

相似文献

许多智力努力需要数学问题解决,但是这种技能仍然超出了计算机的功能。为了衡量机器学习模型中的这种能力,我们引入了数学,这是一个新的数据集,其中包含12,500个具有挑战性的竞争数学问题。数学中的每个问题都有一个完整的分步解决方案,可用于教授模型以生成答案推导和解释。为了促进未来的研究并提高数学的准确性,我们还贡献了一个大型的辅助预处理数据集,这有助于教授模型数学基础知识。即使我们能够提高数学的准确性,我们的结果表明,即使使用巨大的变压器模型,精度仍然相对较低。此外,我们发现,如果缩放趋势继续进行,则简单地增加预算和模型参数计数对于实现强大的数学推理将是不切实际的。虽然缩放变压器正在自动求解大多数基于文本的任务,但缩放量当前尚未求解数学。为了在数学问题解决方面有更多的关注,我们可能需要更广泛的研究界的新算法进步。
Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.