Multivariate and machine learning approaches for honey botanical origin authentication using near infrared spectroscopy

Multivariate and machine learning approaches for honey botanical origin authentication using near infrared spectroscopy
复制标题

DOI:
10.1177/0967033518824765
复制
发表时间:
2019-02-01
影响因子:
1.8
通讯作者:
Capolongo, Francesca
Capolongo, Francesca
中科院分区:
化学4区
文献类型:
--
作者:
Bisutti, Vittoria;Merlanti, Roberta;Capolongo, Francesca

文献摘要

被引文献

相似文献

在这项工作中,近红外光谱的可行性进行了评估,结合化学计量学方法,作为一种工具,119蜂蜜样品的植物来源预测。四个品种有关的polyfloral,金合欢,板栗,和林登树,其特点是首先通过他们的物理化学参数,然后分析了一式三份,使用近红外分光光度计配备了光路黄金反射器。三种不同的分类器建立在不同的多元和机器学习方法蜂蜜植物分类。偏最小二乘判别分析被用作第一种方法来建立蜂蜜分类的预测模型。光谱预处理命名为autoscale,标准正态变量,去趋势,一阶导数和平滑被应用于减少散射的存在下的颗粒大小,如葡萄糖晶体。偏最小二乘判别分析模型的描述性统计的值允许足够的花组预测的金合欢和polyfloral蜂蜜,但不是在板栗和林登的情况下。第二个分类器,基于支持向量机,允许更好的分类金合欢和polyfloral,也实现了板栗的分类。然而,林登样本仍然没有被分类。进一步的调查,旨在提高植物的歧视,利用了一个名为Boruta的特征选择算法,它分配了一个池的39个信息平均近红外光谱变量上的典型判别分析进行了评估。典型判别分析占一个更好的分离样品的植物来源比偏最小二乘判别分析。所使用的方法已允许实现一个完整的认证的金合欢蜂蜜,但不是一个精确的分离的多花的。在预测中重要的变量和Boruta池之间的比较表明,提供信息的波长部分共享,特别是在近红外光谱范围的中波段和远波段。
In this work the feasibility of near infrared spectroscopy was evaluated combined with chemometric approaches, as a tool for the botanical origin prediction of 119 honey samples. Four varieties related to polyfloral, acacia, chestnut, and linden were first characterized by their physical-chemical parameters and then analyzed in triplicate using a near infrared spectrophotometer equipped with an optical path gold reflector. Three different classifiers were built on distinct multivariate and machine learning approaches for honey botanical classification. A partial least squares discriminant analysis was used as a first approach to build a predictive model for honey classification. Spectra pretreatments named autoscale, standard normal variate, detrending, first derivative, and smoothing were applied for the reduction of scattering related to the presence of particle size, like glucose crystals. The values of the descriptive statistics of the partial least squares discriminant analysis model allowed a sufficient floral group prediction for the acacia and polyfloral honeys but not in the cases of chestnut and linden. The second classifier, based on a support vector machine, allowed a better classification of acacia and polyfloral and also achieved the classification of chestnut. The linden samples instead remained unclassified. A further investigation, aimed to improve the botanical discrimination, exploited a feature selection algorithm named Boruta, which assigned a pool of 39 informative averaged near infrared spectral variables on which a canonical discriminant analysis was assessed. The canonical discriminant analysis accounted a better separation of samples according to the botanical origin than the partial least squares discriminant analysis. The approach used has permitted to achieve a complete authentication of the acacia honeys but not a precise segregation of polyfloral ones. The comparison between the variables important in projection and the Boruta pool showed that the informative wavelengths are partially shared especially in the middle and far band of the near infrared spectral range.