Random Forests machine learning applied to gas chromatography - Mass spectrometry derived average mass spectrum data sets for classification and characterisation of essential oils

Random Forests machine learning applied to gas chromatography - Mass spectrometry derived average mass spectrum data sets for classification and characterisation of essential oils
复制标题

DOI:
10.1016/j.talanta.2019.120471
复制
发表时间:
2020-02-01
期刊:
影响因子:
6.1
通讯作者:
Paull, Brett
Paull, Brett
中科院分区:
化学1区
文献类型:
--
作者:
Lebanov, Leo;Tedone, Laura;Paull, Brett

文献摘要

被引文献

相似文献

各种精油(EO)的化学特征的差异来自于每个植物物种和化学型具有独特的次生代谢的事实。因此,这些差异可作为环氧乙烷分类和质量判定的化学标志物。本文中,随机森林(RF)机器学习算法被应用于20种不同的EO的分类。从三路原始气相色谱-质谱数据,总色谱平均质谱(TCAMS)和段平均质谱(SAMS)创建。通过对整个色谱图上各混合物的响应求平均值生成TCAMS,通过对色谱图内某个时间段内各片段的响应求平均值生成SAMS。将RF模型应用于两个数据集,并通过评价预处理数据、树数和每个节点分割中使用的变量数进行优化。通过交叉验证过程评估模型的性能,通过将整个样本集分为训练和验证子集重复50次。超过50个不同训练TCAMS数据集的计算平均袋外误差(OOBE)为3.22 +/-1.29%,而SAMS为2.28 +/-1.33%。通过嵌套交叉验证过程确定EO分类所需的最小变量数。每个步骤中减少的变量的量为10%。结果表明,6个变量的TCAMS数据集与30个变量的SAMS数据集具有相似的预测能力。TCAMS和SAMS对20例EOs分类的OOBE分别为2.89 +/- 1.44%和3.70 +/-1.73%。样品之间的接近度被用来评估它们的质量。具有较大类内接近度的样品具有良好的相似性,而较低的样品则表明化学特征的变化较大。与TCAMS相比,SAMS数据集显示出上级的质量保证潜力。
Differences in chemical profiles of various essential oils (EOs) come from the fact that each plant species and chemotype has a distinctive secondary metabolism. Therefore, these differences can be used as the chemical markers for EO classification and determination of their quality. Herein, the Random Forests (RF) machine learning algorithm was applied to the classification of 20 different EOs. From three-way raw gas chromatography - mass spectra data, total chromatogram average mass spectra (TCAMS) and segment average mass spectra (SAMS) were created. TCAMS was generated by averaging response of each mix over the whole chromatogram and SAMS by averaging the response of each fragment across a certain time segment within the chromatogram. The RF model was applied to the two data sets and optimised through the evaluation of preprocessed data, number of trees, and number of variables used in each node split. The performance of the model was evaluated through a cross-validation process, repeated 50 times by dividing the whole sample set into training and validation subsets. The calculated average out-of-bag error (OOBE), over 50 different training TCAMS data sets was 3.22 +/- 1.29%, while for SAMS it was found to be 2.28 +/- 1.33%. The minimal number of variables necessary for EO classification was determined by a nested cross-validation process. The amount of reduced variables in each step was 10%. It was shown that the TCAMS data set with 6 variables had similar prediction power as the SAMS with 30 variables. OOBE for classification of 20 EOs was 2.89 +/- 1.44% and 3.70 +/- 1.73%, for TCAMS and SAMS, respectively. Proximity between samples was used to evaluate their qualities. Samples with greater intra-class proximity had good similarity, while the lower ones indicated greater variations in the chemical profiles. The SAMS data set showed superior potential for quality assurance, compared with TCAMS.