Measuring the Relative Similarity and Difficulty Between AI Benchmark Problems

Measuring the Relative Similarity and Difficulty Between AI Benchmark Problems
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security
影响因子:
--
通讯作者:
Christopher Pereyda;L. Holder
Christopher Pereyda;L. Holder
中科院分区:
其他
文献类型:
--
作者:
Christopher Pereyda;L. Holder

文献摘要

相似文献

用于评估人工智能系统的挑战性问题、算法测试和数据集激增。然而,没有任何方法可以客观地衡量这些问题的共同困难或它们的相似之处。这是创建更通用的AI系统的障碍。我们提出了一个理论来衡量成对问题之间的相似性。我们通过利用基于深度神经网络的方法来评估这一理论,以客观地测量使用基础数据集的测试问题之间的这些属性。这些方法的实现,然后被用来衡量众所周知的数据集之间的相似性。结果表明,该方法成功地识别了问题之间的相似性和困难性。这可以用来确保用于评估AI系统的测试套件的多样性。
There has been an explosion of challenge problems, algorithmic tests and datasets for evaluating AI systems. Yet no methodology exists to objectively measure either the collective difficulty of these problems or their similarity. This is an obstacle to creating more general AI systems. We pro-pose a theory for measuring the similarity between pair-wise problems. We evaluate this theory by utilizing a methodology based on a deep neural network to objectively measure these properties between test problems using foundational datasets. An implementation of these methods is then used to measure the similarity between well known datasets. Results show that the proposed measure successfully identifies the difficulty and similarity among problems. This can be used to ensure diversity in test suites used to evaluate AI systems.