Measuring the Relative Similarity and Difficulty Between AI Benchmark Problems
Measuring the Relative Similarity and Difficulty Between AI Benchmark Problems
复制标题
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Christopher Pereyda;L. Holder
中科院分区:
文献类型:
--
作者:
Christopher Pereyda;L. Holder
There has been an explosion of challenge problems, algorithmic tests and datasets for evaluating AI systems. Yet no methodology exists to objectively measure either the collective difficulty of these problems or their similarity. This is an obstacle to creating more general AI systems. We pro-pose a theory for measuring the similarity between pair-wise problems. We evaluate this theory by utilizing a methodology based on a deep neural network to objectively measure these properties between test problems using foundational datasets. An implementation of these methods is then used to measure the similarity between well known datasets. Results show that the proposed measure successfully identifies the difficulty and similarity among problems. This can be used to ensure diversity in test suites used to evaluate AI systems.