Genetic algorithm-based method for selecting wavelengths and model size for use with partial least-squares regression: application to near-infrared spectroscopy.

Genetic algorithm-based method for selecting wavelengths and model size for use with partial least-squares regression: application to near-infrared spectroscopy.
复制标题

DOI:
10.1021/ac9607121
复制
发表时间:
1996-12
影响因子:
7.4
通讯作者:
Arjun S. Bangalore;Ronald E. Shaffer;Gary W. Small;Mark A. Arnold
Arjun S. Bangalore;Ronald E. Shaffer;Gary W. Small;Mark A. Arnold
中科院分区:
化学1区
文献类型:
--
作者:
Arjun S. Bangalore;Ronald E. Shaffer;Gary W. Small;Mark A. Arnold

文献摘要

被引文献

相似文献

遗传算法(GAs)被用来实现一个自动化的波长选择程序,用于建立基于偏最小二乘回归的多元校正模型。该方法还允许在构建校准模型时使用的潜变量的数量随着波长的选择而沿着被优化。用于测试这种方法的数据来自通过近红外光谱法测定含水有机物。采用的三个数据集集中于测定(1)1-160 ppm范围内的水中甲基异丁基酮,(2)含有牛血清白蛋白和三醋精的磷酸盐缓冲液基质中葡萄糖的生理水平,以及(3)人血清基质中葡萄糖。这些数据集的特点是分析物信号接近检测限,并存在显着的光谱干扰。研究了光谱数据的信号和噪声特征,并通过实验设计技术为每个数据集找到了遗传算法的最佳配置。尽管光谱数据的复杂性,GA程序被发现执行良好,导致校准模型,显着优于那些基于全光谱分析。此外,显着减少建立模型所需的光谱点的数量被实现。
Genetic algorithms (GAs) are used to implement an automated wavelength selection procedure for use in building multivariate calibration models based on partial least-squares regression. The method also allows the number of latent variables used in constructing the calibration models to be optimized along with the selection of the wavelengths. The data used to test this methodology are derived from the determination of aqueous organic species by near-infrared spectroscopy. The three data sets employed focus on the determination of (1) methyl isobutyl ketone in water over the range of 1-160 ppm, (2) physiological levels of glucose in a phosphate buffer matrix containing bovine serum albumin and triacetin, and (3) glucose in a human serum matrix. These data sets feature analyte signals near the limit of detection and the presence of significant spectral interferences. Studies are performed to characterize the signal and noise characteristics of the spectral data, and optimal configurations for the GA are found for each data set through experimental design techniques. Despite the complexity of the spectral data, the GA procedure is found to perform well, leading to calibration models that significantly outperform those based on full spectrum analyses. In addition, a significant reduction in the number of spectral points required to build the models is realized.