A unified classifiability analysis framework based on meta-learner and its application in spectroscopic profiling data

A unified classifiability analysis framework based on meta-learner and its application in spectroscopic profiling data
复制标题

基于元学习器的统一分类分析框架及其在光谱分析数据中的应用

DOI:
10.1007/s10489-021-02810-8
复制
发表时间:
2021
影响因子:
5.3
通讯作者:
Wang Haiyan
Wang Haiyan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhang Yinsheng;Zhang Zhengyong;Wang Haiyan

文献摘要

相似文献

光谱分析数据(例如,拉曼光谱和质谱)与机器学习相结合,为区分任务提供了一种数据驱动的方法。在这些任务中,研究人员通常从简单的分类模型开始。如果一个模型不起作用,他们将尝试更复杂的模型。如果所有模型都失败了,研究人员将认为数据集是“不可分割的”。这种“试错法”的实践揭示了一个基本问题:数据集是否具有当前判别任务所必需的统计能力?这种“可分类性分析”是数据驱动管道中一个隐含的、经常被忽视的步骤。本文旨在设计一个统一的可分类性分析方法框架。在该框架中,元学习者模型结合了多样化的原子度量(例如,贝叶斯误差率/不可约误差、分类准确度、信息增益/互信息)转化为一个统一的度量(d)。我们已经成功地使用所提出的框架来分析光谱分析数据集,以区分不同年龄的年份酒。5年与16年白酒差异显著(d= 1.447.d> 0.8表示差异显著)。
Spectroscopic profiling data (e.g., Raman spectroscopy and mass spectroscopy), combined with machine learning, have provided a data-driven approach for discriminative tasks. In these tasks, researchers often start with simple classification models. If one model doesn’t work, they will try more sophisticated models. If all models fail, the researchers will deem the data set as “inseparable.“ This “trial-and-error” practice reveals a fundamental question: does the dataset possess the necessary statistical power for the current discriminative task? This “classifiability analysis” is an implicit and often neglected step in the data-driven pipeline. This paper aims to design a unified methodological framework for classifiability analysis. In this framework, a meta-learner model combines diversified atom metrics (e.g., Bayes error rate / irreducible error, classification accuracy, information gain / mutual information) into one unified metric (d). We have successfully used the proposed framework to analyze a spectroscopic profiling dataset to discriminate vintage liquors of different ages. A significant difference (d= 1.447.d> 0.8 indicates a significant difference) between 5-year and 16-year liquors.