Data Augmentation for Mathematical Objects

Data Augmentation for Mathematical Objects
复制标题

DOI:
10.48550/arxiv.2307.06984
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Tereso Del Rio Almajano;M. England
Tereso Del Rio Almajano;M. England
中科院分区:
其他
文献类型:
--
作者:
Tereso Del Rio Almajano;M. England

文献摘要

相似文献

本文讨论并评估了数学对象背景下的数据平衡和数据增强的想法:符号计算和可满足性检查社区在利用机器学习技术来优化其工具时的一个重要主题。我们考虑非线性多项式问题的数据集以及为圆柱代数分解选择变量排序来解决这些问题的问题。通过交换已标记问题中的变量名称,我们生成新的问题实例,当将选择视为分类问题时,不需要任何进一步的标记。我们发现这种增强使 ML 模型的准确性平均提高了 63%。我们研究了这种改进的哪些部分是由于数据集的平衡以及由于进一步增加数据集的大小而实现的,得出的结论是两者都有非常显着的效果。我们通过反思如何将这一想法应用于数学中机器学习的其他用途来完成本文。
This paper discusses and evaluates ideas of data balancing and data augmentation in the context of mathematical objects: an important topic for both the symbolic computation and satisfiability checking communities, when they are making use of machine learning techniques to optimise their tools. We consider a dataset of non-linear polynomial problems and the problem of selecting a variable ordering for cylindrical algebraic decomposition to tackle these with. By swapping the variable names in already labelled problems, we generate new problem instances that do not require any further labelling when viewing the selection as a classification problem. We find this augmentation increases the accuracy of ML models by 63% on average. We study what part of this improvement is due to the balancing of the dataset and what is achieved thanks to further increasing the size of the dataset, concluding that both have a very significant effect. We finish the paper by reflecting on how this idea could be applied in other uses of machine learning in mathematics.