Item Response Theory for Evaluating Regression Algorithms

Item Response Theory for Evaluating Regression Algorithms
复制标题

用于评估回归算法的项目响应理论

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Joint Conference on Neural Network
影响因子:
--
通讯作者:
Telmo de Menezes e Silva Filho
Telmo de Menezes e Silva Filho
中科院分区:
--
文献类型:
--
作者:
João V. C. Moraes;Jessica T. S. Reinaldo;R. Prudêncio;Telmo de Menezes e Silva Filho

文献摘要

被引文献

相似文献

项目反应理论(IRT)是心理测量学中的一种工具,它根据人类对不同难度项目的反应来测量被试的潜在能力。最近,项目反应理论被应用于人工智能的评估,将算法视为应答者,将人工智能任务视为项目。特别是在机器学习中,IRT已经被应用于基于分类器对每个测试实例的预测来评估分类器。基于响应矩阵(分类器与实例),IRT模型估计每个实例的潜在难度和区分度,以及每个分类器的能力,当分类器倾向于正确分类最困难的实例时,它会获得高能力值。以前用于分类评估的IRT模型不直接应用于回归,因为它们依赖于二分响应(即,响应必须是正确的或不正确的)。在本文中,我们提出了一个新的IRT模型,特别是设计用于处理非负无界响应,这是足够的回归算法的绝对误差建模。在该模型中,响应遵循伽马分布,根据受访者的能力和项目的难度和歧视参数参数。与IRT中广泛采用的逻辑曲线相比,所提出的参数化结果使项目特征曲线具有更灵活的形状。提出的模型进行了评估与不同的回归算法和两个基准数据集,一个合成和一个真实的。通过检查这些数据集中呈现不同难度和歧视程度的区域,获得了有用的见解。
Item Response Theory (IRT) is a tool developed in psychometrics to measure latent abilities of human respondents based on their responses to items with different levels of difficulty. Recently, IRT has been applied to evaluation in AI, by treating the algorithms as respondents and the AI tasks as items. Particularly in machine learning, IRT has been applied for evaluation of classifiers based on their predictions to each test instance. Based on a matrix of responses (classifiers vs instances), the IRT model estimates the latent difficulty and discrimination of each instance, as well as the ability of each classifier, in such a way that a classifier receives high ability value when it tends to correctly classify the most difficult instances. The IRT models previously adopted for evaluation in classification are not directly applied for regression, since they rely on dichotomous responses (i.e., a response has to be either correct or incorrect). In this paper we propose a new IRT model, particularly designed for dealing with nonnegative unbounded responses, which is adequate for modelling the absolute errors of regression algorithms. In the proposed model, responses follow a gamma distribution, parameterised according to respondents’ abilities and items’ difficulty and discrimination parameters. The proposed parameterisation results in item characteristic curves with more flexible shapes compared to the logistic curves widely adopted in IRT. The proposed model was evaluated with diverse regression algorithms and two benchmark datasets, one synthetic and one real. Useful insights were derived by inspecting regions in these datasets that present different levels of difficulty and discrimination.