Sigma-z random forest, classification and confidence

Sigma-z random forest, classification and confidence
复制标题

DOI:
10.1088/1361-6501/aaf466
复制
发表时间:
2018-12
影响因子:
2.4
通讯作者:
Alberto Fornaser;M. Cecco;P. Bosetti;Teruhiro Mizumoto;K. Yasumoto
Alberto Fornaser;M. Cecco;P. Bosetti;Teruhiro Mizumoto;K. Yasumoto
中科院分区:
工程技术3区
文献类型:
--
作者:
Alberto Fornaser;M. Cecco;P. Bosetti;Teruhiro Mizumoto;K. Yasumoto

文献摘要

被引文献

相似文献

机器学习是一个很有前途的研究课题,最近取得了显著的成果,导致了更传统的方法被自动学习的解决方案所取代。最近的一些工作已经开始强调,机器学习的模型如何通过对数据应用微小的变化来欺骗,从而导致完全错误的结果。这种行为可以追溯到两个因素:缺乏对传递给模型的投入的任何计量特征,如数据的不确定性,以及缺乏对结果可靠性的评估。本文针对这两个因素,考虑随机森林模型的情况,提出了一种估计置信度的方法,作为分类可靠性的估计器。这考虑了原始分类结构,保持其不变,以及训练数据集的分布。覆盖结构在统计上结合了这两者,并且还在该过程中包括特征不确定性的传播,作为从输入测量得出的另一个元素。新的分类结果是概率向量,该概率向量定义独立于其他类别地将特征条目分配给所考虑的每个类别的可靠性。在这个新的结构中,一个额外的分类结果自然变得可用:不可分类的特征条目。
Machine learning is a promising research topic that has recently achieved remarkable results, leading to the substitution of more traditional methods with automatically learned solutions. Some recent works have begun to highlight how a machine-learned model can be tricked by just applying small variations to the data, resulting in completely erroneous outcomes. Such behaviour can be traced to two elements: the lack of any metrological characterization of the inputs passed to the model, such as the uncertainty of the data, and the lack of an assessment of the reliability of the results. This paper tackles both these elements, considering the case of random forest model and proposing a method for assessing a confidence probability as an estimator for classification reliability. This considers the original classification structure, leaving it untouched, and the distribution of the training datasets. An overlaying structure statistically combines the two, and also includes in the process the propagation of feature uncertainties as a further element deriving from input measurements. The new classification outcome is a vector of probabilities that define how reliably a feature entry can be assigned, or not, to each of the considered classes, independently of others. In this new structure, an additional classification result naturally becomes available: the unclassifiable feature entry.