Evaluating pointwise reliability of machine learning prediction

Evaluating pointwise reliability of machine learning prediction
复制标题

DOI:
10.1016/j.jbi.2022.103996
复制
发表时间:
2022-01-24
影响因子:
4.5
通讯作者:
Bellazzi, Riccardo
Bellazzi, Riccardo
中科院分区:
医学3区
文献类型:
--
作者:
Nicora, Giovanna;Rios, Miguel;Bellazzi, Riccardo

文献摘要

被引文献

相似文献

人们对机器学习应用程序解决临床和生物学问题的兴趣与日俱增。这是由许多研究论文中报告的有希望的结果、基于人工智能的软件产品数量的增加以及对人工智能解决复杂问题的普遍兴趣推动的。因此,提高机器学习输出的质量并增加保障措施以支持其采用非常重要。除了监管和后勤策略之外,一个重要方面是检测机器学习模型何时无法泛化到新的未见过的实例,这些实例可能源自与训练群体相距较远的群体或来自代表性不足的亚群体。因此,机器学习模型对这些实例的预测可能经常是错误的,因为该模型是在其“可靠”工作空间之外应用的,从而导致最终用户(例如临床医生)的信任度降低。因此,当模型在实践中部署时,当模型的预测可能不可靠时向用户提供建议非常重要,特别是在高风险应用程序中,包括医疗保健领域。然而,每个机器学习预测的可靠性评估仍然没有得到很好的解决。在这里,我们回顾了可以支持识别不可靠预测的方法,我们协调了相关概念的符号和术语,并且我们强调和扩展了概念之间可能的相互关系和重叠。然后,我们根据 ICU 院内死亡预测的模拟和真实数据证明了一个可能的综合框架,用于识别可靠和不可靠的预测。为此,我们提出的方法实施了两个互补原则,即密度原则和局部拟合原则。密度原理验证我们要评估的实例与训练集相似。局部拟合原则验证训练模型在与评估实例更相似的训练子集上表现良好。我们的工作可以有助于巩固机器学习领域的工作,尤其是医学领域的工作。
Interest in Machine Learning applications to tackle clinical and biological problems is increasing. This is driven by promising results reported in many research papers, the increasing number of AI-based software products, and by the general interest in Artificial Intelligence to solve complex problems. It is therefore of importance to improve the quality of machine learning output and add safeguards to support their adoption. In addition to regulatory and logistical strategies, a crucial aspect is to detect when a Machine Learning model is not able to generalize to new unseen instances, which may originate from a population distant to that of the training population or from an under-represented subpopulation. As a result, the prediction of the machine learning model for these instances may be often wrong, given that the model is applied outside its "reliable" space of work, leading to a decreasing trust of the final users, such as clinicians. For this reason, when a model is deployed in practice, it would be important to advise users when the model's predictions may be unreliable, especially in high-stakes applications, including those in healthcare. Yet, reliability assessment of each machine learning prediction is still poorly addressed.Here, we review approaches that can support the identification of unreliable predictions, we harmonize the notation and terminology of relevant concepts, and we highlight and extend possible interrelationships and overlap among concepts. We then demonstrate, on simulated and real data for ICU in-hospital death prediction, a possible integrative framework for the identification of reliable and unreliable predictions. To do so, our proposed approach implements two complementary principles, namely the density principle and the local fit principle. The density principle verifies that the instance we want to evaluate is similar to the training set. The local fit principle verifies that the trained model performs well on training subsets that are more similar to the instance under evaluation. Our work can contribute to consolidating work in machine learning especially in medicine.